Thinking Machines Unveils TML-Interaction-Small, a Real-Time Multimodal AI for Conversational Interactivity
Thinking Machines Lab has announced its groundbreaking new multimodal AI system, TML-Interaction-Small, which promises to revolutionize real-time conversational AI. The model is engineered to process diverse inputs—audio, video, and text—simultaneously, breaking down the conventional barriers between input completion and output generation. This concurrent processing capability is a significant departure from many existing multimodal systems that typically operate in a turn-based fashion, requiring users to finish their input before the AI responds.
The core innovation behind TML-Interaction-Small lies in its 'encoder-free early fusion' architecture. Instead of relying on large, pre-trained encoders for each modality, which are common in other multimodal systems like those using OpenAI Whisper for audio or vision transformers for images, Thinking Machines Lab's model directly integrates discretized audio tokens, embeddings of 40x40 pixel image patches (generated by a hierarchical multilayer perceptron), and text embeddings. This integrated approach allows the model to learn connections across modalities more directly from the outset.
The system generates both audio and text outputs via a sophisticated flow-matching decoder. This design enables TML-Interaction-Small to maintain a continuous, responsive dialogue, making interactions feel more natural and immediate. The model achieves this by pairing a fast interaction component for real-time conversation processing with an asynchronous background model dedicated to deeper reasoning. It operates by interleaving 200-millisecond chunks of input processing and output generation, a method the lab refers to as 'micro-turns.'
With 276 billion parameters, TML-Interaction-Small is noted as the largest model of its kind specifically trained for interactive performance among those with publicly disclosed sizes. This substantial scale contributes to its ability to handle complex, real-time scenarios. The potential applications for such a responsive AI are vast, including critical areas like athletic coaching, where immediate feedback is crucial, or monitoring surgical procedures, where real-time analysis can be life-saving.
Thinking Machines Lab, founded approximately 15 months ago by Mira Murati, is currently conducting tests on TML-Interaction-Small, with plans to make it publicly available later in 2026. This announcement places Thinking Machines Lab among a growing number of companies, including OpenBMB, Google, Alibaba, and OpenAI, that are also developing and launching real-time multimodal AI models, signaling a rapid advancement in the field of interactive artificial intelligence.
Read original source