TechniqueRLHF / Alignment8 recent entries12 Aug 2026Pair-Centric Graph Rewiring for Over-Squashing via Optimal Transport-Guided Communication AlignmentarXiv:2608.10619v1 Announce Type: new Abstract: Message-passing neural networks (MPNNs) often struggle when task-relevant information is distributed across distant regions of a graph, since local prop→12 Aug 2026Invertible Logits Transformation for Accuracy-Preserving Post-Hoc Uncertainty CalibrationarXiv:2608.10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining. An ideal calibrator should correct nonl
TechniqueRAG8 recent entries7 Aug 2026A Six-Dimensional Taxonomy of Post-Training Adaptation Techniques with Applications in AI GovernancearXiv:2608.06246v1 Announce Type: new Abstract: Post-training adaptation has become central to modern machine learning practice and includes techniques such as retraining, fine-tuning, parameter-effic→10 Aug 2026Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided CertificatesarXiv:2608.06762v1 Announce Type: new Abstract: Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair an→11 Aug 2026What Would Fix This RAG Failure? Auditing Counterfactual Response with Paired Evidence InterventionsarXiv:2608.08944v1 Announce Type: cross Abstract: A failed retrieval-augmented generation (RAG) answer can be consistent with several unseen responses to evidence repair. We introduce Pair-ID, an offl→11 Aug 2026PreGress: Ranking-Native Pre-training and Prompting for Graph Node RankingarXiv:2608.09016v1 Announce Type: cross Abstract: Node ranking is a fundamental problem in graph information retrieval, measuring the relative importance of nodes and supporting a wide range of applic→11 Aug 2026Mind the Hook: Source-Level Auditing of Privacy Defenses in Retrieval-Augmented GenerationarXiv:2608.09001v1 Announce Type: cross Abstract: Black-box privacy scores for retrieval-augmented generation (RAG) are difficult to interpret unless the audited defense's active pipeline hook is know→12 Aug 2026REATS: LLM Reasoning-based Ensemble Learning for Adaptive Time Series ForecastingarXiv:2608.10149v1 Announce Type: new Abstract: Due to the diversity of real-world time series, no single forecasting model consistently dominates across all samples. Ensemble learning addresses this →12 Aug 2026Post-Calibration Reliability Reranking of Relevance Decisions via Label-wise Monotone ProjectionarXiv:2608.10406v1 Announce Type: cross Abstract: Web search, product search, and question-answering retrieval systems often assign a relevance label and confidence score to each query-candidate pair.→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction
TechniqueAgents8 recent entries12 Aug 2026Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting LogsarXiv:2608.10050v1 Announce Type: new Abstract: Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business ch→12 Aug 2026MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at ScalearXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured →12 Aug 2026FlowScout: From Execution Feedback to Reliable Tool-Using Agent WorkflowsarXiv:2608.10039v1 Announce Type: new Abstract: Agentic workflows have become an important abstraction for building reliable LLM-based automation systems by organizing large language models (LLMs), to→12 Aug 2026Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous DrivingarXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod→12 Aug 2026Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxyarXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server →12 Aug 2026Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief SystemsarXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach →12 Aug 2026An adaptive and evolvable deep reinforcement learning framework for weather predictionarXiv:2608.09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times. Rather than building yet another architecture, we reframe the for→12 Aug 2026AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecastingarXiv:2608.09959v1 Announce Type: cross Abstract: AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to phy
TechniqueFine-tuning8 recent entries12 Aug 2026MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at ScalearXiv:2608.10333v1 Announce Type: new Abstract: LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured →12 Aug 2026Link-adaptive digital twin for robust physical-layer modeling in hybrid-amplified ultra-wideband optical networksarXiv:2608.10517v1 Announce Type: cross Abstract: Accurate physical-layer modeling is increasingly essential for reliable ultra-wideband operation and capacity optimization, especially under the inten→12 Aug 2026GLAM: Efficient Continual Learning at Scale via Grouped LoRA Adapter MergingarXiv:2509.13211v4 Announce Type: replace Abstract: The ability to learn continuously over time remains a major challenge for modern machine learning systems, even in the era of Foundation Models. Whi→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction→12 Aug 2026DEFT: Data-Efficient Frequency-domain Top-k Sampling via Inverse Discrete Fourier Transform for Spatiotemporal Dynamical Systems ModelingarXiv:2608.11019v1 Announce Type: new Abstract: Modeling spatiotemporal dynamical systems governed by partial differential equations (PDEs) poses two major challenges: it either requires expensive phy→12 Aug 2026Can Bayesian Optimization Efficiently Find a Strong Single Expert in Neural Thickets?arXiv:2608.10867v1 Announce Type: new Abstract: Gradient-free post-training has emerged as a compelling alternative to gradient-based optimization for large language models (LLMs), but existing approa→12 Aug 2026Benchmarking Time Series Generation Methods for Privacy-Preserving ForecastingarXiv:2608.10891v1 Announce Type: new Abstract: Time series forecasting in privacy-sensitive domains often requires training models on released data rather than original observations. Synthetic time s→12 Aug 2026A Systematic Sample Size Analysis of ML-Based Path Loss Prediction for LPWANarXiv:2608.11083v1 Announce Type: cross Abstract: Low Power Wide Area Networks like LoRa are increasingly deployed for smart city applications, requiring accurate path loss prediction for effective ne
TechniqueMultimodal8 recent entries11 Aug 2026Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous VehiclesarXiv:2608.08815v1 Announce Type: new Abstract: Traffic sign recognition (TSR) models based on deep neural networks achieve strong clean-data performance but remain vulnerable to physically realizable→11 Aug 2026Deep Multimodal Wearable Sensor Fusion for Detection of Body-Focused Repetitive BehaviorsarXiv:2608.09830v1 Announce Type: new Abstract: Body-focused repetitive behaviors, such as hair pulling and skin picking, are compulsive motor actions commonly associated with obsessive-compulsive and→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation PoliciesarXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f→11 Aug 2026CONFER: Conflict-Aware Evidence Negotiation for Regime-Calibrated Weak Supervision in Multimodal Emotion RecognitionarXiv:2608.07867v1 Announce Type: new Abstract: Multimodal emotion recognition often treats self-reported labels as reliable supervision while overlooking self-report unreliability and cross-modal con→11 Aug 2026Closing the loop in learning with missing dataarXiv:2608.09030v1 Announce Type: cross Abstract: What should a machine learning model learn when data is missing during training? We look at the learning process from a dynamical systems perspective,→11 Aug 2026Auditing Instruction-Trajectory Mismatches in Multimodal Robot DemonstrationsarXiv:2608.07895v1 Announce Type: cross Abstract: Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behavi→12 Aug 2026ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy DistillationarXiv:2608.10905v1 Announce Type: new Abstract: On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Exi→12 Aug 2026Deciding When to Switch: E-Processes for Adaptive Minimax Training for Generative Adversarial NetsarXiv:2608.10096v1 Announce Type: cross Abstract: Modern data science increasingly gives rise to hypothesis-testing problems that are not naturally formulated in terms of parameters within prespecifie
TechniqueSafety8 recent entries12 Aug 2026Dreamer-SAC: Off-Policy Learning in Latent World Models for Sample-Efficient Autonomous DrivingarXiv:2608.10386v1 Announce Type: new Abstract: Sample-efficient reinforcement learning for autonomous driving is often limited by the trade-off between data efficiency and model bias. While world mod→12 Aug 2026Do Time-Series Forecasters Use the Right History: Recoverability, Recovery, and Functional Use of Temporal DelaysarXiv:2608.10433v1 Announce Type: new Abstract: Forecast accuracy does not tell us which past inputs produced a prediction. We separate three questions for time-series models with known delay structur→12 Aug 2026Diffract: Spectral View of LLM Domain AdaptationarXiv:2608.10850v1 Announce Type: new Abstract: We study continual pre-training (CPT) as a mechanism for adapting general-purpose large language models to specialized domains: mathematics, instruction→12 Aug 2026CRHT: A Continuous Regression Hybrid Transformer for Vessel Trajectory Prediction with Online Cluster SamplingarXiv:2608.10256v1 Announce Type: new Abstract: Accurate vessel trajectory prediction is critical for maritime safety and anomaly detection, yet existing models often struggle with geographic bias and→12 Aug 2026Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex OptimizationarXiv:2310.15976v4 Announce Type: replace Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients. Several standard analyses →12 Aug 2026Boundary-Seeking Policy Gradient for Safe Reinforcement LearningarXiv:2608.10204v1 Announce Type: new Abstract: Safe reinforcement learning maximizes reward subject to safety constraints. For Constrained Markov Decision Processes, the linear-programming view over →12 Aug 2026Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxyarXiv:2608.10532v1 Announce Type: cross Abstract: Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server →12 Aug 2026A Joint-Distribution Route to Fair Representations with Continuous Sensitive AttributesarXiv:2608.10470v1 Announce Type: new Abstract: Fair representation learning with a continuous sensitive attribute S requires a representation Z that is statistically independent of S. Existing criter