Agents
I guess image input is the big capability of the models, and tool use can be a substitute for non-omni model output. Still, multimodal voice…
Ethan Mollick notes that image input represents the primary advanced capability of current AI models, and that tool‑use can effectively replace outputs from non‑omni models. He observes that multimoda
Ethan Mollick notes that image input represents the primary advanced capability of current AI models, and that tool‑use can effectively replace outputs from non‑omni models. He observes that multimodal voice integration remains largely untapped, with OpenAI being one of the few developers actively pursuing it.
Related
- Thinking While Speaking: Inference-Time Knowledge Transfer for Responsive and Intelligent Conversational Voice Agents
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- RadAgent: A tool-using AI agent for stepwise interpretation of chest computed tomography
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
Source: Ethan Mollick (X) | 2026-07-13