Have you ever tried to put your face into an AI-generated image, only for it to look like a flat sticker slapped onto a new background? This is a common issue known as the 'copy-paste artifact', and it's a major limitation for most text-to-image AI.
Researchers from Fudan University and StepFun decided to tackle this problem head-on. In their paper, "WithAnyone", they reveal that the core issue is a lack of proper training data. To fix this, they built a massive new dataset with 2 million images, containing hundreds of diverse photos for thousands of different individuals. This allows their new AI model, WithAnyone, to learn the true 'essence' of a person's identity, rather than just copying a single photo.
The results are a significant leap forward for controllable, identity-consistent image generation. WithAnyone can create high-quality images that preserve a person's identity across different poses, expressions, and lighting conditions, effectively breaking the old trade-off between looking accurate and looking natural. This research paves the way for a future of more realistic and expressive AI-generated portraits and group photos.
Cited paper:
H. Xu et al. (2025). WithAnyone: Towards Controllable and ID Consistent Image Generation. arXiv:2510.14975v1. http://arxiv.org/abs/2510.14975v1
Images shown are page renders from the paper PDF for commentary/education.