TechniqueRLHF / Alignment8 recent entries12 Aug 2026ELMER: Evolutionary Language Model that Explores and RefinesarXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a→12 Aug 2026Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction HistoriesarXiv:2608.10319v1 Announce Type: cross Abstract: Large language model (LLM)-powered agents have rapidly evolved from code-completion tools into solvers of complex software engineering tasks. As devel
TechniqueRAG8 recent entries12 Aug 2026Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop EmbeddingarXiv:2608.10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding contro→12 Aug 2026MIRA: Medical Image Reflection for Agentic DiagnosisarXiv:2608.10827v1 Announce Type: cross Abstract: Medical visual agents can use tools to inspect images and retrieve external knowledge, but indiscriminate tool use may introduce noisy or misleading e→12 Aug 2026MAP-Graph: Provenance-Aware Shared Memory for Multi-Agent WorkflowsarXiv:2608.10509v1 Announce Type: new Abstract: Shared memory helps language-model agents reuse information across long workflows, yet relevant evidence may not be admissible for a particular agent or→12 Aug 2026Graphical Models of False Information and Fact Checking EcosystemsarXiv:2208.11582v2 Announce Type: replace-cross Abstract: The wide spread of false information online, including misinformation and disinformation, has become a major problem for our highly digitised →12 Aug 2026GraFine: Retrieval-Time Refinement for Efficient Graph RAG over Corpus GraphsarXiv:2601.18579v2 Announce Type: replace-cross Abstract: Graph RAG on corpus graphs enhances retrieval by leveraging intermediate node content as contextual clues to uncover unretrieved oracle nodes.→12 Aug 2026EvoMem: Memory-Augmented Evolution for Code OptimizationarXiv:2608.10795v1 Announce Type: new Abstract: Successful mutation strategies in evolutionary code search may contain reusable knowledge that is useful beyond a single run, and in some cases may tran→12 Aug 2026Covert Visual Prompt Injection against Commercial Multimodal Large Language ModelsarXiv:2603.29418v2 Announce Type: replace-cross Abstract: Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior l→12 Aug 2026ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and GirlsarXiv:2608.11200v1 Announce Type: cross Abstract: Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or
TechniqueAgents8 recent entries12 Aug 2026ChemWorld: Programmable Chemical Worlds for Controlled and Replayable Agent ExperimentationarXiv:2608.10792v1 Announce Type: new Abstract: Autonomous chemistry increasingly depends on environments in which agents can repeatedly act, observe, and adapt.Physical laboratories provide essential→12 Aug 2026Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn InteractionarXiv:2608.10239v1 Announce Type: new Abstract: Generative AI makes social-engineering attacks more fluent, adaptive, and scalable, increasing the need for LLM-based de- fenders that can protect users→12 Aug 2026Bandwidth-Efficient Multi-Agent Communication through Information Bottleneck and Vector QuantizationarXiv:2602.02035v2 Announce Type: replace-cross Abstract: Multi-agent reinforcement learning systems deployed in real-world robotics applications face severe communication constraints that significant→12 Aug 2026Automating and Scaling Behavioral Scientific Research on AI AgentsarXiv:2608.10030v1 Announce Type: new Abstract: As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI→12 Aug 2026Agentic Instruction Data Selection: Let DataMaster Interpret Your IntentarXiv:2608.10579v1 Announce Type: new Abstract: Although existing instruction data selection methods have introduced various metrics, the inherent complexity of real-world datasets makes it impractica→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions→12 Aug 2026A HamNoSys-Guided Dataset and Baselines for Fine-Grained Isolated Handshape Recognition in Sign LanguagearXiv:2608.10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically →12 Aug 2026A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona ProblemarXiv:2608.10760v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) has become the de-facto interface for connecting LLM agents to enterprise tools, and adoption has been explosive: wit
TechniqueFine-tuning8 recent entries12 Aug 2026Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-TuningarXiv:2608.10473v1 Announce Type: cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction→12 Aug 2026CHORUS: Complementary Experts for High-Coverage Testbench Stimulus GenerationarXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation al→12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu→12 Aug 2026Can Computational Reducibility Lead to Transferable Models for Graph Combinatorial Optimization?arXiv:2603.02462v2 Announce Type: replace-cross Abstract: A key challenge in developing unified neural solvers for combinatorial optimization (CO) is the efficient generalization of models from a give→12 Aug 2026BPG: Balancing Plasticity and Generalization for Domain Incremental LearningarXiv:2608.10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradat→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions→12 Aug 2026A Fixed-Point Neural Operator for Size- and Functional-Transferable Hamiltonian PredictionarXiv:2606.14498v2 Announce Type: replace-cross Abstract: Predicting the Kohn-Sham Hamiltonian with machine learning can accelerate density functional theory while retaining access to molecular orbita→12 Aug 2026A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language ModelsarXiv:2608.10939v1 Announce Type: cross Abstract: Multilingual short-text classification supports operational systems such as content moderation, customer support routing, and intent recognition, yet
TechniqueMultimodal8 recent entries12 Aug 2026DIMOS: Disentangling Instance-level Moving Object SegmentationarXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, an→12 Aug 2026DashArena: Benchmarking LLMs on Interactive Analytic Dashboard GenerationarXiv:2608.10567v1 Announce Type: new Abstract: Analytic dashboards combine coordinated views and interactions for data exploration and decision-making. Recent models can generate them from data and n→12 Aug 2026Covert Visual Prompt Injection against Commercial Multimodal Large Language ModelsarXiv:2603.29418v2 Announce Type: replace-cross Abstract: Although multimodal large language models (MLLMs) are increasingly deployed in real-world applications, their instruction-following behavior l→12 Aug 2026Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and ReliancearXiv:2608.10434v1 Announce Type: new Abstract: Machine learning-based Intrusion Detection Systems (IDS) have demonstrated superior performance in securing Unmanned Aerial Vehicle (UAV) networks. Howe→12 Aug 2026Conversational Orchestration for Organic 6GarXiv:2608.10714v1 Announce Type: cross Abstract: The Organic 6G vision of a network of networks spanning an edge-cloud continuum complemented by non-terrestrial resources requires, to realize its pro→12 Aug 2026ComBodied Agents: a New Paradigm of Human-Centric Agentic AIarXiv:2608.10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither ex→12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu→12 Aug 2026BPG: Balancing Plasticity and Generalization for Domain Incremental LearningarXiv:2608.10804v1 Announce Type: cross Abstract: Deep neural networks excel in various tasks but struggle to generalize across evolving data distributions, leading to significant performance degradat
TechniqueSafety8 recent entries12 Aug 2026ELMER: Evolutionary Language Model that Explores and RefinesarXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a→12 Aug 2026Do AI weather models miss extremes?arXiv:2608.09972v1 Announce Type: cross Abstract: First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression →12 Aug 2026DIMOS: Disentangling Instance-level Moving Object SegmentationarXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, an→12 Aug 2026Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-TuningarXiv:2608.10473v1 Announce Type: cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction→12 Aug 2026ComBodied Agents: a New Paradigm of Human-Centric Agentic AIarXiv:2608.10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither ex→12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu→12 Aug 2026Beyond Forecasting: Recasting Volatility Control as a Routing ProblemarXiv:2608.10375v1 Announce Type: cross Abstract: Volatility control converts risk estimates into portfolio exposure, yet existing approaches often rely on a fixed volatility estimator or a pre-define→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions