Research

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

arXiv:2608.13385v1 Announce Type: new Abstract: Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned tra

DGX agentpaper
researcharxiv-cs-cv

arXiv:2608.13385v1 Announce Type: new Abstract: Implicit multimodal in-context learning compresses demonstrations into internal interventions, ranging from static task vectors to query-conditioned transformations and attention routing. Despite their common goal, these methods differ substantially in how the intervention depends on the query and where it modifies the model, leaving unclear which additional complexity is necessary for a given task. We propose the Selection--Realization Hypothesis. It views demonstrations as inducing a compact family of internal changes from which the query selects, while the model's computation constrains how the selected change can be implemented. We evaluate this account using controlled multimodal tasks in which query dependence varies without changing the underlying task primitives or prompt format. By contrasting correct demonstrations with matched counterfactuals, we measure the structure of explicit M-ICL and test whether it predicts intervention behavior. We find that the success of a static task vector is closely tied to how much of the demonstration-induced change is shared across queries. Additional intervention complexity becomes useful when explicit M-ICL contains query-specific or distributed structure that a local additive shift cannot recover. These relationships extend to natural VQA benchmarks and support cost-aware method selection without access to test performance. Our results provide a unified empirical theory of when demonstrations can be compressed into a task vector and when a more expressive intervention is warranted.

Related

Source: arXiv cs.CV | 2026-08-14

Loading related sources…