CompanyAnthropic8 recent entries28 Jul 2026AlloBench: Measuring Online Tool Allocation Capability in LLM AgentsarXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer→30 Jul 2026Shared SFT Lessons Across Alignment, Model Organisms, and Toy ModelsarXiv:2607.26173v1 Announce Type: new Abstract: Alignment training, model organisms, and toy models are usually treated as separate research areas. But projects in all three frequently use supervised
CompanyOpenAI8 recent entries28 May 2026Inversely Learning Transferable Rewards via Abstracted StatesarXiv:2501.01669v4 Announce Type: replace Abstract: Inverse reinforcement learning (IRL) has progressed significantly toward accurately learning the underlying rewards in both discrete and continuous →2 Jun 2026Lagrangian Perturbation Diffusion Steering: Latent Reinforcement Learning for Generative PoliciesarXiv:2606.01151v1 Announce Type: new Abstract: Behavior cloning with high-capacity generative policies achieves strong imitation performance, but is often limited by demonstration coverage and distri→3 Jun 2026CoughSense: Five-Class Respiratory Disease Classification via Whisper Encoder Fine-Tuning and Dual-Encoder Cross-Attention Fusion with Balanced Contrastive LearningarXiv:2606.02998v1 Announce Type: new Abstract: Automated cough analysis offers a path to low-cost respiratory screening, but most existing work stops at binary COVID-19 detection. A practical tool ne→11 Jun 2026Last-Iterate Convergence of Optimistic Multiplicative Weight UpdatearXiv:2606.11773v1 Announce Type: cross Abstract: Optimistic Gradient Descent Ascent (OGDA) and Optimistic Multiplicative-Weights Update (OMWU) are two very popular algorithms to solve convex/concave →23 Jun 2026Fast-TurboQuant: A Multiplier-Free Online Vector Quantization ApproacharXiv:2606.21448v1 Announce Type: new Abstract: As large language models scale, memory bandwidth for key-value caches and retrieval-augmented generation systems becomes a critical bottleneck. While 1-→4 Aug 2026From Information to Delegation: Mapping Human-AI Financial Decision MakingarXiv:2608.02100v1 Announce Type: cross Abstract: As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become →7 Aug 2026LLM Inference Under Bursty Workload Distribution: Modifying the WAIT AlgorithmarXiv:2608.06135v1 Announce Type: new Abstract: Large Language Models (LLMs) such as ChatGPT and Claude are widely used for information retrieval and problem-solving. Recent work has focused on improv→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation PoliciesarXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f
CompanyGoogle8 recent entries31 Jul 2026Representation and Invariance in Reinforcement LearningarXiv:2112.07752v4 Announce Type: replace-cross Abstract: Researchers have formalized reinforcement learning (RL) in different ways. If an agent in one RL framework is to run within another RL framewo→4 Aug 2026Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent SafetyarXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc→4 Aug 2026The Condition-Number Barrier in Sparse Least SquaresarXiv:2608.02588v1 Announce Type: cross Abstract: In [AS21], Axiotis and Sviridenko conjectured that the linear dependence on the restricted condition number in sparse convex optimization cannot be im→4 Aug 2026Real-Time Detection and Repair of LLM Agent FailuresarXiv:2608.02464v1 Announce Type: cross Abstract: LLM agents fail mid-episode -- they loop, cascade tool errors, drift off goal, fabricate results, or silently absorb corrupted content -- and the stan→4 Aug 2026From Information to Delegation: Mapping Human-AI Financial Decision MakingarXiv:2608.02100v1 Announce Type: cross Abstract: As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become →7 Aug 2026Observation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous ControlarXiv:2608.05989v1 Announce Type: new Abstract: Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning →7 Aug 2026Clinician input steers AI toward accurate and harmful recommendationsarXiv:2603.14158v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are entering clinical workflows, yet evaluations rarely assess how clinician reasoning shapes model behavior duri→11 Aug 2026RouteGuard: Certifying Routing Gain in LLM Multi-Agent Systems When Complementarity Is Not EnougharXiv:2608.07583v1 Announce Type: cross Abstract: Multi-agent LLM systems route among model-backed advisors, yet a deployer rarely knows before shipping whether routing will help at all. Prevailing ro
CompanyMeta8 recent entries10 Aug 2026The Sparsity WhispererarXiv:2608.06630v1 Announce Type: new Abstract: Pruning reduces the inference cost of large language models, but existing criteria primarily preserve large activations or reconstruct layer outputs. We→11 Aug 2026When Skills Meet Safety: Benchmarking and Characterizing the Adaptive Jailbreak Robustness of Skill-Merged LLMsarXiv:2608.08542v1 Announce Type: new Abstract: Model merging has become the default way to give an aligned language model new skills without retraining: a practitioner folds task vectors from math, c→11 Aug 2026When Do Task Vectors Interfere? Mapping the Validity Boundaries of Weight-Space CompositionarXiv:2608.09490v1 Announce Type: new Abstract: Task arithmetic treats fine-tuning displacements as composable directions in weight space, yet it remains unclear when parameter addition reflects predi→11 Aug 2026SPECTRA: Pushing the KV Cache Beyond the 2-Bit Cliff via Spectral Transform CodingarXiv:2608.07915v1 Announce Type: new Abstract: Large language models (LLMs) increasingly read long inputs in the agentic era, from whole documents and codebases to conversations across many turns. Th→11 Aug 2026Gradient Under Microscope: Benchmarking Resource Utilization of Memory-Efficient Gradient Computation MethodsarXiv:2608.08961v1 Announce Type: new Abstract: AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training o→11 Aug 2026Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-TrainingarXiv:2608.08224v1 Announce Type: new Abstract: Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed insi→12 Aug 2026Physics-informed Diffusion Generative Model for Time-Series Data Synthesis in Dynamic SystemsarXiv:2608.10941v1 Announce Type: new Abstract: Industrial time-series signals, such as turbine temperature and rotational speed in aero-engines, are essential for monitoring the health and operationa→12 Aug 2026Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief SystemsarXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach
CompanyMistral8 recent entries4 Aug 2026RAP: KV-Cache Compression via RoPE-Aligned PruningarXiv:2602.02599v4 Announce Type: replace Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and compute of the key-value (KV) cache. Structured pruning is →4 Aug 2026Geometric Analysis of Token Selection in Multi-Head AttentionarXiv:2602.01893v2 Announce Type: replace-cross Abstract: We present a geometric framework for analysing multi-head attention in large language models (LLMs). Without altering the mechanism, we view s→4 Aug 2026Feed-Forward Steering in Transformer Residual DynamicsarXiv:2608.02071v1 Announce Type: new Abstract: Attention-only dynamical theories model Transformer residual directions as particles aggregating on a sphere. We extend this framework by incorporating →4 Aug 2026Efficiency vs. Alignment: Investigating Safety and Fairness Risks in Parameter-Efficient Fine-Tuning of LLMsarXiv:2511.00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks. Prior work sho→5 Aug 2026SAKI: Score-Aware Low-Rank Key Indexing for Long-Context KV RetrievalarXiv:2608.03228v1 Announce Type: new Abstract: Existing low rank KV cache methods preserve either model weights or key variance, neither of which directly reflects the attention scores used during in→5 Aug 2026Quantization Effects on Biomedical LLM ReliabilityarXiv:2608.03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, ver→6 Aug 2026From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language ModelsarXiv:2608.04200v1 Announce Type: cross Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically→7 Aug 2026Can Open-Weight LLMs Produce Kernel-Verified Coq Proofs? A Pilot StudyarXiv:2608.05420v1 Announce Type: cross Abstract: Large language models (LLMs) can generate text that resembles a mathematical proof, but resemblance does not establish correctness. A formal proof che
CompanyxAI8 recent entries24 Jul 2026Explanation-Based Runtime Verification for Trustworthy ML-driven Optical NetworksarXiv:2607.20675v1 Announce Type: new Abstract: Machine learning (ML) models are increasingly integrated into optical network automation frameworks to support tasks such as failure management, perform→27 Jul 2026Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based ExplainabilityarXiv:2607.22428v1 Announce Type: cross Abstract: Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug mode→27 Jul 2026CEL: Comprehensive Counterfactual Explanations Library and BenchmarkarXiv:2607.22045v1 Announce Type: new Abstract: Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes w→28 Jul 2026Disentangling Acoustic Cues in Alzheimer's Pathology and Perception: The Roles of Language and GenderarXiv:2607.23977v1 Announce Type: cross Abstract: Acoustic biomarkers show promise for detecting Alzheimer's Disease (AD), yet whether the cues driving diagnostic AI align with those salient to human →28 Jul 2026Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG ClassificationarXiv:2607.24035v1 Announce Type: cross Abstract: Explainable AI (XAI) is used to assess whether artificial intelligence models rely on meaningful patterns, yet explanations that appear plausible for →3 Aug 2026A Novel XAI-Enhanced Quantum Adversarial Networks for Velocity Dispersion Modeling in MaNGA GalaxiesarXiv:2510.24598v2 Announce Type: replace Abstract: Current quantum machine learning approaches often face challenges balancing predictive accuracy, robustness, and interpretability. To address this, →4 Aug 2026Paris as a 15-Minute City: An Explainable AI PerspectivearXiv:2608.00815v1 Announce Type: new Abstract: The 15-minute city promotes access to everyday services within a short walk or bicycle ride, but its relationship with observed mobility remains difficu→12 Aug 2026BREAD: Baseline-Referenced Explanations for Anomaly DiagnosisarXiv:2608.10587v1 Announce Type: new Abstract: Artificial Intelligence (AI)-based prospective anomaly detection methods are increasingly deployed in high-dimensional and nonlinear settings. Among the
CompanyDeepSeek8 recent entries31 Jul 2026Kalman Meets Curriculum: Efficient Dynamic Prompt Selection for Adaptive RL FinetuningarXiv:2607.27610v1 Announce Type: new Abstract: Reinforcement learning (RL) finetuning significantly enhances the reasoning capabilities of large language models (LLMs), yet its effectiveness critical→31 Jul 2026From Expert Reduction to Behavioral Divergence: Tracing Numerical State through Sparse MoE InferencearXiv:2607.28097v1 Announce Type: new Abstract: Mathematically equivalent expert-reduction orders can produce observably different sparse-MoE executions. We isolate this effect in native DeepSeek-V4-F→31 Jul 2026Beyond Geometric Complementarity: Coherent Overlap in Sparse Mixture-of-Experts RoutingarXiv:2607.28308v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected exper→4 Aug 2026Why Formal Monitors Fail: Attack Distribution Entropy as a Coverage Bound for LTL-Based LLM Agent SafetyarXiv:2608.01388v1 Announce Type: cross Abstract: Runtime safety monitors based on Linear Temporal Logic (LTL) and finite automata (FSA) are increasingly deployed to intercept unsafe tool-call sequenc→4 Aug 2026Emergence Invariance: From Symbolized Thought to Interface RefinementarXiv:2608.01548v1 Announce Type: cross Abstract: Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition. Large lan→7 Aug 2026Can Open-Weight LLMs Produce Kernel-Verified Coq Proofs? A Pilot StudyarXiv:2608.05420v1 Announce Type: cross Abstract: Large language models (LLMs) can generate text that resembles a mathematical proof, but resemblance does not establish correctness. A formal proof che→11 Aug 2026When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM EvaluationarXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex→11 Aug 2026Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-ExpertsarXiv:2608.08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study wheth
CompanyNVIDIA8 recent entries5 Aug 2026Accelerating Dynamic Graph Clustering on GPU Architectures with cuGrapharXiv:2608.03695v1 Announce Type: cross Abstract: This work addresses community detection in temporal networks through GPU-accelerated extensions of spectral clustering and modularity-based algorithms→6 Aug 2026SparseDitto: Customizing GPU Kernels for Different Sparsity Patterns with LLM-Based Agentic SystemarXiv:2608.05033v1 Announce Type: cross Abstract: Sparse matrix kernels are fundamental to scientific computing, graph analytics, and machine learning. Their GPU performance depends strongly on the in→6 Aug 2026An Explainable LLM Agent Layer for Open-World Anomaly Detection in Oil WellsarXiv:2608.04041v1 Announce Type: new Abstract: Open-World Learning (OWL) pipelines for oil well anomaly detection have recently been shown to combine autoencoder-based detection, multiclass classific→7 Aug 2026Operating Multi-Node Full Fine-Tuning on NVIDIA B300: A Field Report on Telemetry-Based Triage, Negative Results, and Operational HardeningarXiv:2608.05944v1 Announce Type: cross Abstract: We report operational experience full-fine-tuning a 32.76B-parameter dense model (Qwen3-32B) on 16 x NVIDIA B300 (two nodes, FSDP / ZeRO-3) -- among t→10 Aug 2026SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding PredictionarXiv:2608.06441v1 Announce Type: new Abstract: Full-graph GNN training delivers high accuracy but scales poorly on multi-server clusters due to heavy, irregular inter-node embedding exchanges. We pre→10 Aug 2026Dual-Node NVIDIA DGX Spark over Tailscale: A Remote-Access Testbed for Distributed LLM Training and Cyber-Threat-Intelligence Fine-TuningarXiv:2608.07226v1 Announce Type: cross Abstract: Compact AI systems make local language-model experimentation increasingly accessible, yet practical evidence for multi-node training on desktop-class →11 Aug 2026RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space AttentionarXiv:2608.08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneo→11 Aug 2026FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal ApplicationsarXiv:2607.18171v2 Announce Type: replace Abstract: Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose effici