Hardware

MaineCoon is the first video model that focuses on social interactions: facial expressions, emotions, fluid conversation, audio-lip sync, et…

MaineCoon is the first video model that focuses on social interactions: facial expressions, emotions, fluid conversation, audio-lip sync, etc. Really impressive inference specs: 22B params, 47.5 FPS o

DGX agentx-post
hardwarefrancois-chollet--x

MaineCoon is the first video model that focuses on social interactions: facial expressions, emotions, fluid conversation, audio-lip sync, etc. Really impressive inference specs: 22B params, 47.5 FPS on a single H100. Generates in real-time at <$0.001/sec. They achieve this with an agentic streaming inference framework with 3 different auxiliary models to manage the cache and lookahead buffer. Super cool work. Most AI video today is still: prompt → wait → watch a clip. MaineCoon is built for something different: prompt → talk → interact in real time. In our vision, the character is not a fixed video clip that just waits for your input. It keeps generating voice, expression, and motion …

Source: Francois Chollet (X) | 2026-06-21

Loading related sources…