TechniqueRLHF / Alignment8 recent entries12 Aug 2026Mapping and Measuring the Behavioral Evolution of Large Language ModelsarXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across genera→12 Aug 2026Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text ModelsarXiv:2601.17387v3 Announce Type: replace Abstract: Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether
TechniqueRAG8 recent entries11 Aug 2026DS@GT ARC at Touche: Large Language Models for Retrieval-Augmented DebatearXiv:2608.08143v1 Announce Type: cross Abstract: We extend the DS@GT ARC working-note submission to the Touche 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utt→11 Aug 2026AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document RetrievalarXiv:2608.08732v1 Announce Type: cross Abstract: Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds→11 Aug 2026An Agentic Generative Large Language Model for Treatment Planning of Colorectal CancerarXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure gui→12 Aug 2026The Signal Rail: A Deterministic Motion Grammar for Communicating Conversational Agent State in Terminal InterfacesarXiv:2608.10689v1 Announce Type: cross Abstract: Terminal interfaces to conversational agents report rich internal state (listening, thinking, executing tools, awaiting input, failing) almost entirel→12 Aug 2026TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language ModelsarXiv:2509.25143v2 Announce Type: replace-cross Abstract: Existing medical reasoning benchmarks for vision-language models primarily focus on analyzing a patient's condition based on an image from a s→12 Aug 2026Self-Knowledge Retrieval Augmented Generation Framework for Patent MatchingarXiv:2608.11030v1 Announce Type: cross Abstract: Patent retrieval and matching based on large language models (LLMs) play a vital role in intellectual property protection. However, due to the complex→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMsarXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud→12 Aug 2026Order Matters: LVLMs as Judges for Temporal Reasoning in Image SequencesarXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud
TechniqueAgents8 recent entries12 Aug 2026Mitigating Context Interference for Reliable and Efficient Search AgentsarXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are s→12 Aug 2026InSight-doc: Agentic Visual Perception for Long-Document UnderstandingarXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we→12 Aug 2026Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and BiasesarXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NL→12 Aug 2026DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student DistillationarXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri→12 Aug 2026Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition AgentsarXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen→12 Aug 2026Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human DesignarXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, s→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent HarnessesarXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these →12 Aug 2026Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using AgentsarXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa
TechniqueFine-tuning8 recent entries12 Aug 2026Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASRarXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-reso→12 Aug 2026Reinforcement Learning-based Semi-supervised Knowledge Distillation with LLM-as-a-JudgearXiv:2604.02621v2 Announce Type: replace Abstract: Reinforcement Learning (RL) substantially improves the reasoning capabilities of language models, but most existing RL fine-tuning approaches rely e→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMsarXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud→12 Aug 2026myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASRarXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work prese→12 Aug 2026Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-TrainingarXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog FaithfulnessarXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;→12 Aug 2026Data Attribution of Emergent Misalignment with Persona FeaturesarXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leadi→12 Aug 2026Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate ShiftarXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar
TechniqueMultimodal8 recent entries11 Aug 2026AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document RetrievalarXiv:2608.08732v1 Announce Type: cross Abstract: Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds→12 Aug 2026VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world→12 Aug 2026TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language ModelsarXiv:2509.25143v2 Announce Type: replace-cross Abstract: Existing medical reasoning benchmarks for vision-language models primarily focus on analyzing a patient's condition based on an image from a s→12 Aug 2026StreamFlow: Dynamic Memory Flows for Streaming Video UnderstandingarXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under→12 Aug 2026Order Matters: LVLMs as Judges for Temporal Reasoning in Image SequencesarXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud→12 Aug 2026MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level AlignmentarXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image repr→12 Aug 2026FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion EditingarXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc→12 Aug 2026DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student DistillationarXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri
TechniqueSafety8 recent entries12 Aug 2026Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition AgentsarXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen→12 Aug 2026Data Attribution of Emergent Misalignment with Persona FeaturesarXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leadi→12 Aug 2026ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question AnsweringarXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many→12 Aug 2026Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate ShiftarXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar→12 Aug 2026Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus TheoryarXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers d→12 Aug 2026Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent HarnessesarXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these →12 Aug 2026Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online SafetyarXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis →12 Aug 2026Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using AgentsarXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa