Familiar

Real-Time AI Dubbing. Neolab for World Translation.

YC Summer 2026Artificial IntelligenceGenerative AIB2BEntertainment

About Familiar

Audio-visual translation infrastructure for all video, live and recorded -- the first and only that works in real time. World translation: not just changing voice and lips into a new language, but understanding the scene, the sound, and preserving the human. ๐Ÿฆ‹ Movies & TV, Micro-dramas, Live Broadcasts, Live Shopping, Corporate & Education, Ads, Creators & Podcasts, Platforms. Day-and-date, in every territory. Founders An Zhu Liu, Jibin Song, Mingi Kwon, and Xu Zheng lead a team of PhDs, professors, and researchers from the world's top universities who started the field of joint audio-visual human models and co-authored the current state of the art in real-time human animation. Since 2021: over 10,000 citations and 100+ papers at ICLR, NeurIPS, CVPR, ECCV, and more. Translation has to be a world model. A translator has to know who's speaking: where every sound in the dense polyphony originates (laughter, layered crowds). Laughter here, a short yell there, a voice from off-screen, someone on screen mouthing words before being interrupted. Everyone else scaffolds separate models and pre-processing; a world model simulates the scene. It's a spatio-temporal reasoner: predicting continuity across frames forces it to internalize 3D geometry, spatial reasoning, and permanence -- of objects, people, and the environment itself. The scene state survives the edit: when the shot pans or cuts away, someone off-screen still exists, still owns their voice, and is still there when the camera comes back. Multiple people, side views, off angles, dense polyphony, fast motion -- the hard cases hold. Vs ElevenLabs: they lose the laughter, the music, the ambiance -- 270% more background-sound error -- and ours sounds 27.8% more like you. That means no recasting, no ADR, no M&E stems required -- it works from the final mix. Vs human translators: 100.3% of their quality -- scalable, real time. Published benchmarks: https://thefamiliarlab.com/benchmark

Founders

  • Mingi Kwon

    Founder

    AI PhD from Yonsei University. Built large-scale text-to-video models at Adobe Research and published at top AI conferences. Previously operated localized YouTube channels with 400K+ total subscribers.

    LinkedIn โ†—

  • An Zhu Liu

    Founder

    Stanford dropout. Tinkerer & researcher. Built the most realistic real-time human avatars with unmatched cloning accuracy early 2026. Trained personally on multiple nodes of B200s.

    LinkedIn โ†—X โ†—

  • Xu Zheng

    Founder

    AI PhD from HKUST, 30+ publications on top-tier AI venues with 2k+ citations and 2k+ github stars.

    LinkedIn โ†—

Discussion

Posting anonymously

No comments yet โ€” start the discussion.