TechniqueRLHF / Alignment8 recent entries12 Aug 2026Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban ScenesarXiv:2608.10954v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) demonstrate impressive performance in benign scenarios, their cognitive reliability deteriorates signif→12 Aug 2026ENTLORE: A Graph-Grounded Benchmark for Latent Organizational Reasoning in Enterprise Question AnsweringarXiv:2608.10679v1 Announce Type: cross Abstract: Enterprise question answering is framed as retrieving internal documents and generating grounded answers. Routine enterprise records, however, are wor
TechniqueRAG8 recent entries12 Aug 2026Situation Graph Prediction for User Perspective ModelingarXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data →12 Aug 2026Rethinking Text-Based Image Retrieval in Specific DomainarXiv:2608.10524v1 Announce Type: cross Abstract: Driven by the rapid advancement of vision-language representation learning, Text-based Image Retrieval (TBIR) has made notable progress. However, exis→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMsarXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud→12 Aug 2026Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR EvaluationarXiv:2608.10385v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assess→12 Aug 2026MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent WorkflowsarXiv:2608.10509v1 Announce Type: new Abstract: Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or→12 Aug 2026How OneAdvanced deployed over 50 AI agents on UK-sovereign AWSLearn how OneAdvanced, a UK enterprise software provider, built a UK-sovereign AI platform by self-hosting Llama 4 Maverick and Llama Guard 4 on Amazon SageMaker AI, with a RAG pipeline on pgvector an→12 Aug 2026DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous DrivingarXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction
TechniqueAgents8 recent entries12 Aug 2026Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxyarXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server →12 Aug 2026Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief SystemsarXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach →12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent HarnessesarXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these →12 Aug 2026An adaptive and evolvable deep reinforcement learning framework for weather predictionarXiv:2608.09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the for→12 Aug 2026AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecastingarXiv:2608.09959v1 Announce Type: cross Abstract: AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to phy→12 Aug 2026Ahrefs launches AI agent workspace Letaido for marketers and agenciesMarketing intelligence company Ahrefs Pte. Ltd. today launched Letaido, an agent-powered marketing workspace built to take over the recurring research, reporting and monitoring work that fills up a ma→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions→12 Aug 2026A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign LanguagearXiv:2608.10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically
TechniqueFine-tuning8 recent entries12 Aug 2026DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous DrivingarXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog FaithfulnessarXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction→12 Aug 2026DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation PredictionarXiv:2608.10595v1 Announce Type: cross Abstract: Proteolysis-targeting chimeras (PROTACs) induce protein degradation by recruiting a target protein to an E3 ubiquitin ligase, making degradation a joi→12 Aug 2026CHORUS: Complementary Experts for High-Coverage Testbench Stimulus GenerationarXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation al→12 Aug 2026BPG: Balancing Plasticity and Generalization for Domain Incremental LearningarXiv:2608.10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradat→12 Aug 2026Benchmarking Time Series Generation Methods for Privacy-Preserving ForecastingarXiv:2608.10891v1 Announce Type: new Abstract: Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time s→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions
TechniqueMultimodal8 recent entries12 Aug 2026DriveVLA-M0: Failure-Aware Memory Augmentation for Autonomous DrivingarXiv:2608.10413v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for end-to-end autonomous driving by enabling unified reasoning across→12 Aug 2026DreamOmni3: Scribble-based Editing and GenerationarXiv:2512.22525v2 Announce Type: replace Abstract: Recently unified generation and editing models have achieved remarkable success with their impressive performance. These models rely mainly on text →12 Aug 2026DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student DistillationarXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri→12 Aug 2026Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial NetsarXiv:2608.10096v1 Announce Type: cross Abstract: Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecifie→12 Aug 2026DashArena: Benchmarking LLMs on Interactive Analytic Dashboard GenerationarXiv:2608.10567v1 Announce Type: new Abstract: Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and n→12 Aug 2026ComBodied Agents: a New Paradigm of Human-Centric Agentic AIarXiv:2608.10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither ex→12 Aug 2026CapProbe: Evaluating Detailed Image Captions via Full-Scene Dense Question AnsweringarXiv:2608.11074v1 Announce Type: new Abstract: Evaluating detailed image captions from Vision-Language Models (VLMs) requires going beyond surface-level semantic similarity. Reference-based metrics (→12 Aug 2026BPG: Balancing Plasticity and Generalization for Domain Incremental LearningarXiv:2608.10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradat
TechniqueSafety8 recent entries12 Aug 2026Introspective Attention Modulation for Safe Text-to-Image GenerationarXiv:2607.14945v2 Announce Type: replace Abstract: State-of-the-art flow based text-to-image (T2I) models exhibit remarkable generative abilities but remain vulnerable to producing unsafe content. Pr→12 Aug 2026GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation ReasoningarXiv:2608.10494v1 Announce Type: new Abstract: Earth observation (EO) agents construct scientifically valid tool workflows and ground their conclusions in current geospatial evidence. This is challen→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction→12 Aug 2026ComBodied Agents: a New Paradigm of Human-Centric Agentic AIarXiv:2608.10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither ex→12 Aug 2026Boundary-Seeking Policy Gradient for Safe Reinforcement LearningarXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over →12 Aug 2026Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxyarXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server →12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent HarnessesarXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these →12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions