TechniqueRLHF / Alignment8 recent entries8 Apr 2026Governance-Aware Agent Telemetry for Closed-Loop Enforcement in Multi-Agent AI SystemsEnterprise multi-agent AI systems produce thousands of inter-agent interactions per hour, yet existing observability tools capture these dependencies without enforcing anything. OpenTelemetry and Lang→28 Apr 2026StereoFoley: Object-Aware Stereo Audio Generation from VideoWe present StereoFoley, a video-to-audio generation framework that produces semantically aligned, temporally synchronized, and spatially accurate stereo sound at 48 kHz. While recent generative video-
TechniqueRAG1 recent entries15 Jul 2026CLaRa: Bridging Retrieval and Generation with Continuous Latent ReasoningRetrieval-augmented generation (RAG) enhances large language models (LLMs) with external knowledge but still suffers from long contexts and disjoint retrieval–generation optimization. In this work, we
TechniqueAgents8 recent entries9 Apr 2026LaCy: What Small Language Models Can and Should Learn is Not Just a Question of LossThis paper was accepted at the Workshop on Memory for LLM-Based Agentic Systems at ICLR. Language models have consistently grown to compress more world knowledge into their parameters, but the knowled→6 May 2026From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMsTrue spatial intelligence for multimodal agents transcends low-level geometric perception, evolving from knowing where things are to understanding what they are for. While existing benchmarks, such as→2 Jul 2026Multi-Agent Teams Hold Experts BackMulti-agent LLM systems are increasingly deployed as autonomous collaborators, where agents interact freely rather than execute fixed, pre-specified workflows. In such settings, effective coordination→7 Jul 2026Weblica: Scalable and Reproducible Training Environments for Visual Web AgentsThe web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories fo→7 Jul 2026FlowEval: Reference-Based Evaluation of Generated User InterfacesWhile large language models (LLMs) and coding agents are often applied to user interface (UI) development, developers find it difficult to reliably assess their proficiency in visual and interaction d→9 Jul 2026Recursive Language Models Meet Uncertainty: The Surprising Effectiveness of Self-Reflective Program Search for Long ContextLong-context handling remains a core challenge for language models: even with extended context windows, models often fail to reliably extract, reason over, and use the information across long contexts→10 Jul 2026Behavioral Privacy Leakage in Agentic Negotiation: Formalizing and Mitigating Inference Attacks via Randomized PoliciesThis paper was accepted at the AI4TCI (Workshop on AI for Secure and Trustworthy Critical Infrastructure Systems) Workshop at the International Conference on Availability, Reliability and Security (AR→14 Jul 2026Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive AssistantsProactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Exi
TechniqueFine-tuning5 recent entries9 Apr 2026LaCy: What Small Language Models Can and Should Learn is Not Just a Question of LossThis paper was accepted at the Workshop on Memory for LLM-Based Agentic Systems at ICLR. Language models have consistently grown to compress more world knowledge into their parameters, but the knowled→23 Jun 2026Metric-Dependent Annotation Saturation for Learning from Label DistributionsWhen annotators disagree on a label, the disagreement itself carries signal—and the number of annotators needed to capture it depends on the evaluation metric. We fine-tune NLI models on label distrib→7 Jul 2026Weblica: Scalable and Reproducible Training Environments for Visual Web AgentsThe web is complex, open-ended, and constantly changing, making it challenging to scale training data for visual web agents. Existing data collection attempts remain limited to offline trajectories fo→14 Jul 2026Multilingual Semantic Retrieval for Apple Music SearchApple Music serves listeners across 150+ storefronts in dozens of languages, with a catalog that grows by hundreds of thousands of new tracks daily. At this scale, search recall on misspelled, transli→6 Aug 2026Locking Pretrained Weights via Deep Low-Rank Residual DistillationThe quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software plat
TechniqueMultimodal8 recent entries6 May 2026From Where Things Are to What They’re For: Benchmarking Spatial–Functional Intelligence for Multimodal LLMsTrue spatial intelligence for multimodal agents transcends low-level geometric perception, evolving from knowing where things are to understanding what they are for. While existing benchmarks, such as→22 May 2026VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant ModelsStreaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time visual assistants. Exis→2 Jul 2026On Robustness and Chain-of-Thought Consistency of RL-Finetuned VLMsReinforcement learning (RL) finetuning has become a key technique for enhancing large language models (LLMs) on reasoning-intensive tasks, motivating its extension to vision language models (VLMs). Wh→9 Jul 2026Incentivizing Temporal-Awareness in Egocentric Video Understanding ModelsMultimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning dep→14 Jul 2026Proactive Agent Research Environment: Simulating Active Users to Evaluate Proactive AssistantsProactive agents that anticipate user needs and autonomously execute tasks hold great promise as digital assistants, yet the lack of realistic user simulation frameworks hinders their development. Exi→15 Jul 2026One Layer Is Enough: Adapting Pretrained Visual Encoders for Image GenerationVisual generative models (e.g., diffusion models) typically operate in compressed latent spaces to balance training efficiency and sample quality. In parallel, there has been growing interest in lever→3 Aug 2026Understanding Alignment in Multimodal LLMs: A Comprehensive StudyPreference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively under→7 Aug 2026Arbitrage: Efficient Reasoning via Advantage-Aware SpeculationModern Large Language Models achieve impressive reasoning capabilities with long Chain of Thoughts, but they incur substantial computational cost during inference, and this motivates techniques to imp
TechniqueSafety8 recent entries8 Apr 2026Governance-Aware Agent Telemetry for Closed-Loop Enforcement in Multi-Agent AI SystemsEnterprise multi-agent AI systems produce thousands of inter-agent interactions per hour, yet existing observability tools capture these dependencies without enforcing anything. OpenTelemetry and Lang→29 Apr 2026DSO: Direct Steering Optimization for Bias MitigationGenerative models are often deployed to make decisions on behalf of users, such as vision-language models (VLMs) identifying which person in a room is a doctor to help visually impaired individuals. Y→8 May 2026RVPO: Risk-Sensitive Alignment via Variance RegularizationCurrent critic-less RLHF methods aggregate multi-objective rewards via an arithmetic mean, leaving them vulnerable to constraint neglect: high-magnitude success in one objective can numerically offset→6 Jul 2026Understanding Annotator Safety Policy with InterpretabilitySafety policies define what constitutes safe and unsafe AI outputs, guiding data annotation and model development. However, annotation disagreement is pervasive and can stem from multiple sources such→7 Jul 2026MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow MatchingRecent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday→9 Jul 2026Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and WhyOn-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this signal is beneficial and under which it is detrimental→9 Jul 2026Incentivizing Temporal-Awareness in Egocentric Video Understanding ModelsMultimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning dep→3 Aug 2026Understanding Alignment in Multimodal LLMs: A Comprehensive StudyPreference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively under