TechniqueRLHF / Alignment8 recent entries11 Aug 2026Gated Spatial Redundancy Projection for Pathology Transformer AttentionsarXiv:2608.08374v1 Announce Type: new Abstract: Transformer models are increasingly used for whole-slide image analysis in computational pathology. Yet, WSIs differ fundamentally from natural images: →11 Aug 2026Diffusion Image Editing via Asynchronous Token DecodingarXiv:2608.09322v1 Announce Type: new Abstract: Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naive
TechniqueRAG8 recent entries11 Aug 2026Temporal Misgrounding in Legal RAG: A Versioned-Corpus Benchmark for French Tax LawarXiv:2608.09393v1 Announce Type: cross Abstract: We identify and quantify temporal misgrounding: the systematic retrieval and citation of the currently in-force version of a legal article when the ap→11 Aug 2026SodaMem: Evidence-Grounded Temporal Graph Memory for LLM AgentsarXiv:2608.08055v1 Announce Type: new Abstract: Large language model (LLM) agents that assist users over weeks of conversation must remember what is currently true, not merely what was once said. Flat→11 Aug 2026Retrieval-Augmented Generation-Based Color Restoration for Low-Light Image EnhancementarXiv:2608.08211v1 Announce Type: cross Abstract: Recent low-light image enhancement (LLIE) methods have driven brightness and structural fidelity close to that of normally-exposed images, yet their o→11 Aug 2026RAG-Based Auto-Configuration for Industrial Fieldbus DevicesarXiv:2608.08618v1 Announce Type: cross Abstract: Industrial device commissioning requires engineers to manually extract hundreds of protocol-specific parameters from heterogeneous PDF manuals and tra→11 Aug 2026Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical LimitsarXiv:2608.07913v1 Announce Type: cross Abstract: Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated→11 Aug 2026DocAtlas: Long-Document Understanding as Mutable-State InteractionarXiv:2608.07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-a→11 Aug 2026AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document RetrievalarXiv:2608.08732v1 Announce Type: cross Abstract: Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds→11 Aug 2026A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual DataarXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment.
TechniqueAgents8 recent entries12 Aug 2026Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model AgentsarXiv:2608.10198v1 Announce Type: new Abstract: Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reaso→12 Aug 2026Nutrition Data Infrastructure for the AI Era: Operationalizing FAIR for Agent-Mediated ResearcharXiv:2608.10363v1 Announce Type: new Abstract: AI agents can accelerate nutrition research, but their analyses inherit the identity, semantic, and release ambiguities of the underlying data. We prese→12 Aug 2026Grok 4.6 is objectively #1 when considering intelligence, speed & costGrok 4.6 is objectively #1 when considering intelligence, speed & cost SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with →12 Aug 2026Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Cla…Grok 4.6 is now one of the top models in the world for agentic workflows It ranks #1 on the Artificial Analysis Agentic Index, tied with Claude Opus 5 Max Agentic AI is about more than answering quest→12 Aug 2026Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code …Grok 4.6 is an excellent model. I’ve been using it heavily for the past couple of weeks and it handles everything from simple coding & code review all the way to designing and debugging complex system→12 Aug 2026give 4.6 a try and let us know how it goes. your feedback is a big part of why the model gets better with each iteration.give 4.6 a try and let us know how it goes. your feedback is a big part of why the model gets better with each iteration. SpaceXAI's Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, j→12 Aug 2026Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware ReformulationarXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie→12 Aug 2026DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student DistillationarXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri
TechniqueFine-tuning8 recent entries10 Aug 2026Meta returns to open weights: Muse Glimmer, its first open-weights release since Llama 4, scores 35 on the Artificial Analysis Intelligence …Meta returns to open weights: Muse Glimmer, its first open-weights release since Llama 4, scores 35 on the Artificial Analysis Intelligence Index. It is a 30B-parameter model, and the first from Meta →10 Aug 2026Index SLM Technical ReportarXiv:2607.09885v2 Announce Type: replace Abstract: We present Index-1.9B, a series of open small language models developed at Bilibili. The series comprises four models: Index-1.9B-Base, a foundation→10 Aug 2026b10342model : Granite-Switch Architecture (#25107) granite-switch: add llama.cpp backend (POC, CPU) New 'granite-switch' architecture: a dense, all-attention Granite-4.1 model with N embedded LoRA adapters →11 Aug 2026ZOMP: Zeroth-Order Multi-Modal Prompt Tuning for Vision-Language ModelsarXiv:2608.08060v1 Announce Type: new Abstract: Fine-tuning vision-language models such as CLIP typically requires backpropagation (BP) through the full model, which is infeasible when only forward-pa→11 Aug 2026Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer m…Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days. Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like a→11 Aug 2026TS-Mob: Social and Geographical-Aware Time Series Foundation-Model Framework for Human Mobility PredictionarXiv:2507.00945v2 Announce Type: replace Abstract: Short-term forecasting of aggregated human mobility flows supports urban planning, intelligent transportation systems, and emergency response, yet e→11 Aug 2026Diffusion Image Editing via Asynchronous Token DecodingarXiv:2608.09322v1 Announce Type: new Abstract: Text-guided diffusion image editing aims to modify semantic attributes of an image while preserving its identity, layout, and background. However, naive→12 Aug 2026Fast and Memory-Efficient Wavelet Convolutions via I/O-Aware ReformulationarXiv:2608.10805v1 Announce Type: cross Abstract: Wavelet convolution (WTConv) has emerged as an increasingly popular drop-in replacement for standard convolutions, expanding a network's receptive fie
TechniqueMultimodal8 recent entries11 Aug 2026RoboSeg: Online Part-Level Semantic Reconstruction for Robotic Manipulation via a Single Eye-in-Hand CameraarXiv:2608.09778v1 Announce Type: new Abstract: Robotic manipulation requires perception systemsthat identify actionable parts such as handles, rims, triggers,and tool tips, not merely object categori→11 Aug 2026Discovering Diverse Planning Policies for Multimodal Embodied Agents with Quality-Diversity OptimizationarXiv:2608.08523v1 Announce Type: new Abstract: Multimodal embodied agents are increasingly required to solve long-horizon tasks by integrating visual observations, textual goals, and interaction hist→11 Aug 2026BRUCE: Benchmarking Robustness Under Corruption Escalation for Scientific Vision-Language ReasoningarXiv:2608.07742v1 Announce Type: new Abstract: Visual-language models (VLMs) frequently struggle with robustness issues in real-world situations due to low- or varying-quality input images. In this p→11 Aug 2026AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document RetrievalarXiv:2608.08732v1 Announce Type: cross Abstract: Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds→11 Aug 2026A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual DataarXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. →12 Aug 2026Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model AgentsarXiv:2608.10198v1 Announce Type: new Abstract: Latent-space communication allows heterogeneous vision-language model agents to exchange continuous representations without serializing visual and reaso→12 Aug 2026DreamOmni3: Scribble-based Editing and GenerationarXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text →12 Aug 2026DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student DistillationarXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri
TechniqueSafety8 recent entries11 Aug 2026Retrieval-Augmented Generation-Based Color Restoration for Low-Light Image EnhancementarXiv:2608.08211v1 Announce Type: cross Abstract: Recent low-light image enhancement (LLIE) methods have driven brightness and structural fidelity close to that of normally-exposed images, yet their o→11 Aug 2026Private Anytime Selective-Risk Certification for Federated Retrieval-Augmented Generation: Guarantees and Empirical LimitsarXiv:2608.07913v1 Announce Type: cross Abstract: Selective-risk certificates promise that accepted outputs meet a declared error target. We develop Fed-SRC, a score-agnostic certificate for federated→11 Aug 2026LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning RoutingarXiv:2608.08889v1 Announce Type: new Abstract: Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large L→11 Aug 2026KumbhDoot: A Scale-Ready, LLM-Bounded Architecture for Mass-Gathering Public-Service AssistantsarXiv:2608.07520v1 Announce Type: cross Abstract: Mass religious gatherings such as the Kumbh Mela concentrate tens of millions of people into a single region over a few weeks, producing intense, repe→11 Aug 2026Knowledge-Distilled End-to-End Reinforcement Learning for Smooth 6-DOF Thrust Control and Rapid Adaptation to Ocean Currents in Remotely Operated VehiclesarXiv:2608.08598v1 Announce Type: new Abstract: With the continuous improvement of computational capabilities, end-to-end reinforcement learning has been rapidly developed for remotely operated vehicl→11 Aug 2026A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual DataarXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. →12 Aug 2026The CASE Framework: A Multi-Disciplinary Control Architecture for Governing Enterprise Agentic AIarXiv:2608.10153v1 Announce Type: new Abstract: Enterprises are deploying autonomous AI agents faster than they can govern them, and prevailing approaches stretch a single discipline, typically DevSec→12 Aug 2026Self-Normalized Inference for Constant-Stepsize Temporal-Difference Learning under Markovian SamplingarXiv:2608.10896v1 Announce Type: cross Abstract: Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account fo