Selected Publications
2026
- CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound EffectsACMMM 2026| ACM International Conference on Multimedia
- Co-Steer: Cross-Modal Collaborative Steering for Jailbreaking MLLMsECCV 2026| European Conference on Computer Vision
- Multi-stage Metric Learning with CLIP-based Adaptation for Few-shot Action RecognitionTMM 2026| IEEE Transactions on Multimedia
- Boot-and-Feedback Framework for Generalist-Expert Model Collaboration in Breast Ultrasound DiagnosisICASSP (oral) 2026| IEEE International Conference on Acoustics, Speech and Signal Processing
- Omni2Sound: A Fundamental Study on Dataset, Base Model, and Benchmark for Unified Video-Text-to-Audio GenerationCVPR 2026| Conference on Computer Vision and Pattern Recognition
- Translating Signals to Languages for sEMG-Based Activity RecognitionCVPR 2026| Conference on Computer Vision and Pattern Recognition
- Fresco: Frequency–Spatial Consistent Optimization for Fine-Grained Head Avatar ModelingCVPR (highlight) 2026| Conference on Computer Vision and Pattern Recognition
2025
- Semantic-guided Cross-Modal Prompt Learning for Skeleton-based Zero-shot Action RecognitionCVPR 2025| Proceedings of the Computer Vision and Pattern Recognition Conference
2024
2023
2022
- Human action recognition from various data modalities: A reviewTPAMI 2022| IEEE Transactions on Pattern Analysis and Machine Intelligence