Model Releases
HounsWorld: A Multimodal World Model for Hidden Patient-State Readout, Reconstruction, and Simulation
arXiv:2608.12904v1 Announce Type: new Abstract: Clinical intelligence requires estimating a patient's underlying condition from incomplete observations rather than learning isolated mappings from scan
arXiv:2608.12904v1 Announce Type: new Abstract: Clinical intelligence requires estimating a patient's underlying condition from incomplete observations rather than learning isolated mappings from scans to answers. Volumetric medical images provide dense observations of anatomy, attenuation, and lesions, whereas clinical language provides sparse but complementary semantic observations. We formulate CT-centered intelligence as inference over a shared latent patient state, under which readout, reconstruction, and simulation all become state-dependent prediction problems. To operationalize this view, we introduce HounsBench, a computed tomography (CT) centric patient-state benchmark that unifies these three task families with patient-disjoint splits and per-family metrics, and HounsWorld, a 3B multimodal world model that treats volumetric scans and language as observations of the shared state through Joint Understanding-Generation Learning. A shared transformer forms an implicit patient-state estimate and supports three outputs: query-conditioned answers that read out the state, reports and captions that reconstruct it in language, and condition-specific CT volumes for low-dose denoising, virtual contrast enhancement, and anatomy-constrained text-and-mask-to-volume generation. Zero-initialized CT adapters preserve pretrained multimodal mappings, while condition-explicit Hounsfield-unit window sampling exposes clinically meaningful density observations. HounsWorld shows strong performance across all three task families while consistently improving CT understanding through clinically structured completion. Our project is available at https://github.com/byhwhite/HounsWorld.git
Related
- Benchmarking Visual State Tracking in Multimodal Video Understanding
- MeniOmni: A Structured Multimodal Benchmark for Holistic Meniscus Injury Assessment
- Exploring the Capabilities of Large Language Model Encoders for Image-Text Retrieval in Chest X-rays
- Fine-tuning a multimodal large language model for clinician-grade autism behavioral scoring from short home videos
- Flex-pi: A Multi-Stream World-Action Model with Compute Flexibility
Source: arXiv cs.CV | 2026-08-14