TechniqueRLHF / Alignment8 recent entries12 Aug 2026Do Time-Series Forecasters Use the Right History: Recoverability, Recovery, and Functional Use of Temporal DelaysarXiv:2608.10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structur→12 Aug 2026DIMOS: Disentangling Instance-level Moving Object SegmentationarXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, an
TechniqueRAG8 recent entries11 Aug 2026An Agentic Generative Large Language Model for Treatment Planning of Colorectal CancerarXiv:2608.09142v1 Announce Type: new Abstract: Treatment planning in precision oncology requires synthesizing heterogeneous patient information with rapidly evolving clinical guidelines to ensure gui→11 Aug 2026Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform forHigh Dose Rate (HDR) BrachytherapyarXiv:2608.08163v1 Announce Type: new Abstract: The convergence of the Metaverse and Large Language Model (LLM)-based AI agent is catalyzing a shift toward autonomous, immersive, and personalized peda→11 Aug 2026A Dynamic-Semantics Framework for Grounding Human Referring Expressions in Visual Perceptual DataarXiv:2608.08663v1 Announce Type: cross Abstract: Humans converge on shared names for novel, hard-to-describe objects through repeated interaction, a process psycholinguists call lexical entrainment. →12 Aug 2026When should you start post-training your own models? @FireworksAI_HQ CEO @lqiao’s answer: after product-market fit. Not because it's hard...…When should you start post-training your own models? @FireworksAI_HQ CEO @lqiao’s answer: after product-market fit. Not because it's hard... but because only after PMF is the data coming off your prod→12 Aug 2026Towards Color-Faithful Low-Light Image Enhancement via Adaptive Color Debiasing and Saturation RectificationarXiv:2608.10512v1 Announce Type: new Abstract: Low-light imaging often introduces color bias caused by the low signal-to-noise ratio and the image formation process. Although recent low-light image e→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMsarXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud→12 Aug 2026Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop EmbeddingarXiv:2608.10207v1 Announce Type: new Abstract: Bus bunching degrades service regularity and increases passenger waiting in high-frequency transit. Existing reinforcement-learning-based holding contro→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction
TechniqueAgents8 recent entries12 Aug 2026ELMER: Evolutionary Language Model that Explores and RefinesarXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a→12 Aug 2026Easy3D-Labels: Supervising Semantic Occupancy Estimation with 3D Pseudo-Labels for Automotive PerceptionarXiv:2509.26087v5 Announce Type: replace Abstract: In perception for automated vehicles, safety is critical not only for the driver but also for other agents in the scene, particularly vulnerable roa→12 Aug 2026Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous DrivingarXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod→12 Aug 2026Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition AgentsarXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen→12 Aug 2026ComBodied Agents: a New Paradigm of Human-Centric Agentic AIarXiv:2608.10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither ex→12 Aug 2026Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxyarXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server →12 Aug 2026// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fi…// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New resea→12 Aug 2026Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using AgentsarXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa
TechniqueFine-tuning8 recent entries12 Aug 2026INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student SimulatorsarXiv:2608.10492v1 Announce Type: new Abstract: Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, w→12 Aug 2026Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-TrainingarXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting→12 Aug 2026ELMER: Evolutionary Language Model that Explores and RefinesarXiv:2608.10196v1 Announce Type: cross Abstract: Program evolution can measure whether a mutation helped, but it rarely controls how far the mutation moves in behavior space. Syntactic edit size is a→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction→12 Aug 2026Data Attribution of Emergent Misalignment with Persona FeaturesarXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leadi→12 Aug 2026Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-TuningarXiv:2608.10473v1 Announce Type: cross Abstract: Offline-to-online (O2O) reinforcement learning aims to leverage policies pretrained on static datasets while improving them through online interaction→12 Aug 2026Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate ShiftarXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar→12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu
TechniqueMultimodal8 recent entries12 Aug 2026LoRCA: LoRA Cycle Adaptation for Histology to HiP-CT Translation with DINOv3arXiv:2608.10002v1 Announce Type: cross Abstract: Hierarchical Phase-Contrast Tomography (HiP-CT) is a synchrotron based X-ray imaging technique that enables non-destructive, volumetric imaging of int→12 Aug 2026Hidden in Plain Sight: Diffusion-Based Unrestricted Robotic Attacks on Vision-Language-Action ModelsarXiv:2608.10393v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have shown strong capabilities in controlling robots across diverse manipulation tasks. However, their adversarial r→12 Aug 2026FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion EditingarXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc→12 Aug 2026DIMOS: Disentangling Instance-level Moving Object SegmentationarXiv:2606.12826v2 Announce Type: replace-cross Abstract: Moving instance segmentation (MIS) attracts increasing attention due to its broad applications in traffic surveillance, autonomous driving, an→12 Aug 2026ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist DeferralarXiv:2608.10885v1 Announce Type: new Abstract: Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imagin→12 Aug 2026ComBodied Agents: a New Paradigm of Human-Centric Agentic AIarXiv:2608.10915v1 Announce Type: new Abstract: After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither ex→12 Aug 2026CARE: Confidence-Aware Reasoning for Reliable Medical VQAarXiv:2608.10964v1 Announce Type: cross Abstract: Reinforcement Fine-Tuning (RFT) has enabled medical Multimodal Large Language Models (MLLMs) to produce Chain-of-Thought (CoT) reasoning for visual qu→12 Aug 2026BooST: Bridging Semantics and Motions for Efficient Skill TransferarXiv:2608.10600v1 Announce Type: cross Abstract: Skill abstraction---the process of learning reusable and temporally extended behaviors---has emerged as a key paradigm for improving sample efficiency
TechniqueSafety8 recent entries12 Aug 2026Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online SafetyarXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis →12 Aug 2026APCReg: Anatomical-Prior-Guided Coarse-to-Fine CBCT--IOS Registration via Multi-View Projection and Reliability-Controlled Residual CorrectionarXiv:2608.09993v1 Announce Type: cross Abstract: Registration between cone-beam computed tomography (CBCT) and intraoral scans (IOS) is essential for patient-specific surgical planning. However, disp→12 Aug 2026AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance LossarXiv:2608.11205v1 Announce Type: new Abstract: Frechet distance has recently emerged as an effective distribution-level objective for generator post-training, complementing the conventional sample-le→12 Aug 2026// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fi…// Actions Speak Louder Than Words // Multilingual agent evaluation compares final answers and throws the trajectory away. The trajectory fixes cost, latency, failure mode, and auditability. New resea→12 Aug 2026Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using AgentsarXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa→12 Aug 2026A Neural Network Based Teleoperation for Remote Controlled VehiclesarXiv:2608.10367v1 Announce Type: new Abstract: Direct teleoperation of vehicles faces critical technical bottlenecks: communication latency and the operator's inability to physically perceive unmodel→12 Aug 2026A Joint-Distribution Route to Fair Representations with Continuous Sensitive AttributesarXiv:2608.10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute S requires a representation Z that is statistically independent of S. Existing criter→12 Aug 2026A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection MethodsarXiv:2608.10203v1 Announce Type: new Abstract: Despite the success of convolutional neural networks in image classification tasks and their general application in multi-modal models, their susceptibi