RT DW A1 Li, Zhangheng T1 Advancing Efficiency and Trustworthiness: From Computer Vision to Multimodal Large Language Models