WeChat Vision, Tencent Inc.
Recent advances in video synthesis have enabled realistic integration of real individuals, driving demand for identity-aware generation. While emerging methods support joint appearance and voice injection in audio-visual models, they primarily focus on single-subject settings. Multimodal identity integration across multiple subjects remains limited, and precise alignment between visual and vocal identities in multi-subject scenarios remains underexplored. We present Identity-as-Presence, a unified framework for joint personalized audio-video generation. An automated data curation pipeline constructs identity-labeled audio-visual pairs for single- and multi-subject scenes. A unified identity injection mechanism then binds paired appearance and voice through shared cross-modal identity binding and subject-anchored captions. A multi-stage training strategy further leverages large-scale unimodal data alongside scarce paired clips to mitigate modality imbalance. Experiments show superior audio quality, video fidelity, and audio-visual consistency, with stronger multi-subject binding than the compared methods.
The overall dual-tower DiT architecture and training framework. The model encodes video, audio, paired appearance/voice references, and subject-anchored captions, binds them with shared identity embeddings, and injects them via reference positioning and decoupled asymmetric attention. Training proceeds in three stages: unimodal identity, joint multimodal alignment, and spatio-temporal multi-view fine-tuning.
@article{chen2026identity,
title={Identity as Presence: Towards Appearance and Voice Personalized Joint Audio-Video Generation},
author={Qin Chen, Yingjie Chen, Shilun Lin, Xing Cai, Binxin Yang, Long Zhou, Qixin Yan, Wenjing Wang, Dingming Liu, Hao Liu, Chen Li, and Jing LYU},
journal={arXiv preprint arXiv:2603.17889},
year={2026}
}