CompanyAnthropic8 recent entries11 Aug 2026A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding AgentsarXiv:2608.09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.→12 Aug 2026VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?arXiv:2608.10408v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world
CompanyOpenAI8 recent entries11 Aug 2026When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply ChainsarXiv:2608.07538v1 Announce Type: new Abstract: As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably→11 Aug 2026Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasksarXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision→11 Aug 2026How to Ask the AI: A User Perspective Survey for Large Language Model PromptingarXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t→11 Aug 2026From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming ExaminationsarXiv:2608.07523v1 Announce Type: cross Abstract: Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language →11 Aug 2026DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?arXiv:2608.07614v1 Announce Type: cross Abstract: Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code→11 Aug 2026Curriculum Generation under Structured Parametric Environments for Robust Navigation PoliciesarXiv:2608.08545v1 Announce Type: cross Abstract: Robust navigation policies for autonomous agents must generalize across continuously varying environmental conditions such as turn rates, obstacles, f→11 Aug 2026An Expectation-Maximization Perspective on Reinforcement Learning for LLM ReasoningarXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective RefinementarXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models
CompanyGoogle8 recent entries11 Aug 2026Can Open-Weight Models Compete on Financial Text Comprehension?arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and MitigationarXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi→11 Aug 2026Automating Deception: Scalable Multi-Turn LLM JailbreaksarXiv:2511.19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves t→11 Aug 2026An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus PhotographyarXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.→11 Aug 2026360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied AgentsarXiv:2608.08814v1 Announce Type: cross Abstract: We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment construc→12 Aug 2026Situation Graph Prediction for User Perspective ModelingarXiv:2602.13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences. Progress is limited by a data →12 Aug 2026Reference-Free Post-Training of Open Large Language Models for Multilingual Machine TranslationarXiv:2608.10812v1 Announce Type: cross Abstract: We study reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised-finetuned MiL→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI AgentsarXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be
CompanyMeta8 recent entries12 Aug 2026When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM ReasoningarXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of→12 Aug 2026Measuring Semantic Abstractness of SAE Features via NonlocalityarXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the c→12 Aug 2026Interpreting Language Model Hidden States at ScalearXiv:2608.10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions d→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog FaithfulnessarXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;→12 Aug 2026Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context ExtensionarXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that th→12 Aug 2026Behavioral Inference at Scale: The Fundamental Asymmetry Between Motivations and Belief SystemsarXiv:2509.05624v3 Announce Type: replace-cross Abstract: How much information about an agent's underlying values can be recovered from its observable behavior? This question matters for any approach →12 Aug 2026Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided SchedulingarXiv:2508.03611v3 Announce Type: replace-cross Abstract: This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions
CompanyMistral8 recent entries11 Aug 2026KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMsarXiv:2608.09779v1 Announce Type: cross Abstract: Answering complex conditional questions using Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) remains a challenge, particularly →11 Aug 2026High-Layer Attention Pruning with RescalingarXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional→11 Aug 2026Do All LLMs Know When They're Being Harmful? A Reproducibility Study of Latent-Space Safety Probes Across Model FamiliesarXiv:2608.08029v1 Announce Type: cross Abstract: Khatri et al. (2026) [DOI: 10.1109/DSN-W70714.2026.00027] show that lightweight MLP probes on final-layer activations of a single 8B model (LLaMA-3.1-→11 Aug 2026DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM InferencearXiv:2608.08878v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequen→11 Aug 2026Can Open-Weight Models Compete on Financial Text Comprehension?arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability→12 Aug 2026The Truth Stays in the Family: Enhancing Contextual Grounding via Inherited Truthful Heads in Model LineagesarXiv:2606.15821v2 Announce Type: replace-cross Abstract: Recent advances in large language models (LLMs) have produced many specialized multimodal LLMs (MLLMs) that share common foundational LLMs, fo→12 Aug 2026REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMsarXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud→12 Aug 2026Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog FaithfulnessarXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;
CompanyxAI8 recent entries4 Aug 2026Does Explainability Transfer? A Controlled Benchmark of Attribution Methods on Vision Transformers and CNNsarXiv:2608.02396v1 Announce Type: new Abstract: Most evidence on the effectiveness of explainable artificial intelligence (XAI) attribution methods has been established on convolutional neural network→5 Aug 2026AI Security Leaderboard: Methodology, Results and Minimal StandardarXiv:2608.03070v1 Announce Type: cross Abstract: Frontier AI model developers increasingly rely on layered safeguards to prevent catastrophic misuse, but little public evidence exists on how much pro→7 Aug 2026AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic SimulationsarXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr→10 Aug 2026Post-Grokking Collapse at the Representation-Readout Interface in Muon-Trained TransformersarXiv:2608.07436v1 Announce Type: new Abstract: Under the standard split, Muon gets hidden matrices and AdamW embeddings/output head. Muon groks modular addition faster, but its solutions do not hold.→10 Aug 2026Evaluating XAI Support From A Hierarchical Reinforcement Learning Policy in Human-Agent CollaborationarXiv:2608.06381v1 Announce Type: cross Abstract: Explainable AI (XAI) has shown promise for human-agent collaboration, yet results rely on hand-crafted policies in custom environments, limiting gener→11 Aug 2026The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical GamesarXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro→11 Aug 2026The Authority Expectancy Effect in Multi-User ConflictarXiv:2608.08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a m→11 Aug 2026Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent MisalignmentarXiv:2608.08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts
CompanyDeepSeek8 recent entries11 Aug 2026Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-ExpertsarXiv:2608.08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study wheth→11 Aug 2026Automated Generation of Complexity-Validated Decision Scenarios Using Large Language ModelsarXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b→11 Aug 2026An Expectation-Maximization Perspective on Reinforcement Learning for LLM ReasoningarXiv:2504.18587v2 Announce Type: replace-cross Abstract: Reinforcement learning has emerged as a powerful approach for improving the reasoning capabilities of large language models, as demonstrated b→11 Aug 2026Adversarial Attacks on Deep OCR SystemsarXiv:2608.07636v1 Announce Type: cross Abstract: Deep-OCR (DeepSeek-OCR) advances document recognition by treating the visual modality as an optical compression medium, enabling long-context OCR at l→12 Aug 2026Rethinking LLM Verification: Evidence Structure, Uncertainty, and Selective RefinementarXiv:2608.10725v1 Announce Type: new Abstract: Large language models (LLMs) often rely on shortcuts rather than systematic reasoning, raising safety concerns in medical applications. Allowing models →12 Aug 2026Persistent Recursive Worlds Enable Autonomous Software EvolutionarXiv:2608.10450v1 Announce Type: cross Abstract: Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve conti→12 Aug 2026Measuring Semantic Abstractness of SAE Features via NonlocalityarXiv:2608.10537v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) have helped uncover mechanistic explanations for LLM behaviours such as reasoning, jailbreaking etc., via understanding the c→12 Aug 2026CHORUS: Complementary Experts for High-Coverage Testbench Stimulus GenerationarXiv:2608.10090v1 Announce Type: new Abstract: Large language models (LLMs) have advanced code generation, where executable feedback provides a more reliable learning signal than textual imitation al
CompanyNVIDIA8 recent entries10 Aug 2026RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMsarXiv:2608.07088v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) encode images as long visual token sequences, making prefilling and KV-cache storage expensive. Existing trai→10 Aug 2026Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning PredictorsarXiv:2608.06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making ac→11 Aug 2026Who Verifies the Benchmark? Decentralizing Trust in Large Language Model EvaluationarXiv:2608.07762v1 Announce Type: new Abstract: LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims tha→11 Aug 2026RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space AttentionarXiv:2608.08081v1 Announce Type: cross Abstract: Large mixture-of-experts (MoE) language models with 26--120 billion parameters exceed the memory capacity of consumer devices through three simultaneo→11 Aug 2026Memorization Dynamics in Knowledge Distillation for Language ModelsarXiv:2601.15394v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improv→11 Aug 2026LGNNIC: Acceleration of Large-Scale GNN Training using SmartNICsarXiv:2608.07733v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) are widely used across domains such as natural sciences, social network analysis, chip design, and recommendation systems→11 Aug 2026Governing the KV Cache: Preventing Timing Side-Channel Leakage in Multi-Tenant LLM InferencearXiv:2608.09225v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary throughput optimization in modern large language model (LLM) inference, enabling prefix reuse across requests.→12 Aug 2026Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR EvaluationarXiv:2608.10385v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assess