XPENG has published TuringViT, a vision encoder for vision-language and vision-language-action models used in smart driving, smart cockpit systems and its IRON humanoid robot program.

XPENG offers TuringViT-18L and TuringViT-24L. At 1536×1536 resolution, the company says TuringViT-18L reached 3.04 times the throughput of Seed1.5-ViT and 2.16 times that of SigLIP2-ViT-L.

XPENG says the model was trained on 850 million image-text pairs and achieved an average score of 83.6% across six zero-shot benchmarks, exceeding open-source baselines trained on 10 billion samples. [IT Home, in Chinese]