CompanyAnthropic8 recent entries11 Aug 2026You Don't Need To Stay in The Loop: An Agentic Robotics Loop for Robot-Policy ImprovementarXiv:2608.07555v1 Announce Type: new Abstract: Coding agents such as Claude Code and Codex close the software loop: a main agent manages the loop, subagents analyze and execute, tools do the work. We→11 Aug 2026When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM EvaluationarXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex
CompanyOpenAI8 recent entries29 Jul 2026AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal EvaluationarXiv:2607.25881v1 Announce Type: new Abstract: We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were indepen→31 Jul 2026LayerRAG-Bench: A Cross-Layer Reliability Benchmark for Agentic Retrieval-Augmented GenerationarXiv:2607.27353v1 Announce Type: new Abstract: Agentic retrieval-augmented generation systems can produce answers that appear grounded while failing at the evidence, tool-contract, authorization, or →31 Jul 2026Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World EnglishesarXiv:2607.28505v1 Announce Type: new Abstract: The rise of generative artificial intelligence (GenAI) in academic writing and publishing (AWP) raises questions about linguistic inclusivity and the le→5 Aug 2026SAT-Edge-Agent: Hardware-in-the-Loop Edge-Agent Orchestration for Onboard Satellite IntelligencearXiv:2608.03728v1 Announce Type: new Abstract: Onboard satellite intelligence requires a task layer that translates mission intent into local tool calls, exposes execution state, and returns machine-→5 Aug 2026ETA: A New Agentic Paradigm for Embodied TasksarXiv:2608.03924v1 Announce Type: new Abstract: When will robots have their ChatGPT moment? Such a breakthrough requires a general-purpose robot that can handle unfamiliar tasks in unfamiliar environm→10 Aug 2026NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge ProofsarXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should→11 Aug 2026Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual ReasoningarXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi→11 Aug 2026How to Ask the AI: A User Perspective Survey for Large Language Model PromptingarXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t
CompanyGoogle8 recent entries7 Aug 2026Vibe Compiler: A Research-Logic Synthesis Tool That Runs without Prompt Engineering -Toward Enhancing Metacognition for Sustaining Agency in the Age of Generative AI-arXiv:2608.05545v1 Announce Type: cross Abstract: Generative AI used as a capable servant has greatly accelerated intellectual work, but it also risks eroding human epistemic agency by encouraging unc→10 Aug 2026Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agentsarXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language mode→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputarXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L→11 Aug 2026LLM within MCP Matters: Measuring Inefficient Resource Utilization Driven by LLMsarXiv:2608.08467v1 Announce Type: new Abstract: The Model Context Protocol (MCP) standardizes how servers expose data and tools to Large Language Models (LLMs). A common server design embeds frequentl→11 Aug 2026Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent MisalignmentarXiv:2608.08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts→11 Aug 2026BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and MitigationarXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi→11 Aug 2026An Agentic AI Framework Overcomes Fundamental Limitations of Large Language Models for Glaucoma Detection from Fundus PhotographyarXiv:2608.07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.→12 Aug 2026Navigation Alone Is Not Enough: Evaluating Explanatory Assistive UI AgentsarXiv:2608.09944v1 Announce Type: cross Abstract: Modern web interfaces are increasingly difficult to use with screen readers, particularly when pages update dynamically or hide important structure be
CompanyMeta8 recent entries10 Aug 2026NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge ProofsarXiv:2608.07167v1 Announce Type: new Abstract: Giving an AI agent the ability to send emails, query databases, or execute commands is useful--until the agent is tricked into doing something it should→11 Aug 2026Persistent Semantic Entities in Tool-Augmented LLM SystemsarXiv:2608.07952v1 Announce Type: cross Abstract: Tool-augmented LLM agents can harbor implicit state that persists across sessions, activates through events, and propagates across agent boundaries---→11 Aug 2026Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent HarnessesarXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harne→12 Aug 2026When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM ReasoningarXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of→12 Aug 2026Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent MemoryarXiv:2608.11066v1 Announce Type: cross Abstract: We prove inference-time quantum coordination advantages for specified AI state-tracking tasks. A solver compresses semantic history into a future-acce→12 Aug 2026MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom GrapharXiv:2608.10504v1 Announce Type: new Abstract: As coding agents increasingly handle implementation, the central challenge shifts from building individual agents to building an infrastructure that sys→12 Aug 2026Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent SystemsarXiv:2608.09964v1 Announce Type: cross Abstract: The Brazilian Conference on Intelligent Systems (BRACIS) is the main national venue for Artificial Intelligence research in Brazil, hosted by the Braz→12 Aug 2026Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic CritiquearXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions
CompanyMistral8 recent entries1 Jun 2026MLIPilot: LLM-Driven Auto-Research for Machine-Learned Interatomic PotentialsarXiv:2605.30889v1 Announce Type: cross Abstract: Constructing production-quality machine-learned interatomic potentials (MLIPs) requires balancing accuracy, dynamical stability, and computational thr→3 Jun 2026TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness AssessmentarXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-w→3 Jun 2026GTBench: A Curriculum-Grounded Benchmark for Evaluating LLMs as Mathematical Research Assistants in Graph TheoryarXiv:2606.03144v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as self-study assistants in technical disciplines, yet their reliability as mathematical reasoning as→3 Jun 2026Generating the Modal Worker: A Cross-Model Audit of Race and Gender in LLM-Generated Personas Across 41 OccupationsarXiv:2510.21011v3 Announce Type: replace-cross Abstract: As generative AI tools are increasingly used to portray people in professional roles, understanding their racial and gender representational b→6 Jun 2026Synapse: Federated Tool Routing via Typed Compendium ArtifactsarXiv:2602.00911v2 Announce Type: replace Abstract: The unit of collaboration in federated learning determines what guarantees are even expressible. Flat units like weights, prompts, raw examples, car→24 Jun 2026Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical AssessmentarXiv:2606.24331v1 Announce Type: new Abstract: Transformer-based language models have become the default substrate for natural language processing and the pace of new releases has made it hard for pr→7 Jul 2026Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety PrimitivesarXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age→11 Aug 2026Measuring the Tokenization Premium: A Cost Audit for Underserved Language CommunitiesarXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe
CompanyxAI8 recent entries23 Jul 2026Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted EscalationarXiv:2607.15434v3 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome→24 Jul 2026Position Bias is Hidden Behind Ceiling Effects: A Permutation Diagnostic for LLM BenchmarksarXiv:2607.20864v1 Announce Type: cross Abstract: Position bias in multiple-choice LLM evaluation is widely cited as a confound in capability comparisons, but published measurements rely on single ans→24 Jul 2026Interpretable Embeddings with Sparse Autoencoders: A Data Analysis ToolkitarXiv:2512.10092v2 Announce Type: replace Abstract: Analyzing large-scale text corpora is a core challenge in machine learning, crucial for tasks like identifying undesirable model behaviors or biases→27 Jul 2026Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based ExplainabilityarXiv:2607.22428v1 Announce Type: cross Abstract: Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug mode→28 Jul 2026ML-based Predictive Models for Power Consumption in Virtualised O-RANsarXiv:2607.24256v1 Announce Type: cross Abstract: As communication networks adopt virtualized and disaggregated architectures, achieving energy efficiency has become increasingly important for both ec→31 Jul 2026AI LEGO: Scaffolding Cross-Functional Collaboration in Industrial Responsible AI Practices during Early Design StagesarXiv:2505.10300v2 Announce Type: replace-cross Abstract: Responsible AI (RAI) efforts increasingly emphasize the importance of addressing potential harms early in the AI development lifecycle through→11 Aug 2026The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical GamesarXiv:2608.09574v1 Announce Type: new Abstract: LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move fro→11 Aug 2026Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent MisalignmentarXiv:2608.08212v1 Announce Type: new Abstract: In-context learning (ICL) can induce emergent misalignment (EM), where narrow misaligned examples alter answers to unrelated questions. Existing prompts
CompanyDeepSeek8 recent entries4 Aug 2026Emergence Invariance: From Symbolized Thought to Interface RefinementarXiv:2608.01548v1 Announce Type: cross Abstract: Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition. Large lan→5 Aug 2026UrbanAgent: A Tool-Augmented Agent for Cross-System Urban TasksarXiv:2608.03018v1 Announce Type: new Abstract: Modern cities rely on an increasing number of digital services to operate, but residents' daily needs are still difficult to meet. Services are fragment→10 Aug 2026Policy-Masked Private Experts: Auditable and Reversible Capability Access Control in Sparse MoE ModelsarXiv:2608.06690v1 Announce Type: cross Abstract: Most language-model access controls regulate behavior while leaving the same computation available to every request. We study a different systems ques→10 Aug 2026Fisher-R1: Training LLM Agents for Reliable Hypothesis TestingarXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t→11 Aug 2026When Grammar Guides the Attack: Uncovering Control-Plane Vulnerabilities in LLMs with Structured OutputarXiv:2503.24191v4 Announce Type: replace-cross Abstract: Content Warning: This paper may contain unsafe or harmful content generated by LLMs that may be offensive to readers. Large Language Models (L→11 Aug 2026When Counterbalancing Hides the Bias: Access-Conditioned Position Lock in Forced-Choice LLM EvaluationarXiv:2607.10202v2 Announce Type: replace Abstract: Forced-choice probes with counterbalanced orientations are a standard tool for measuring language-model 'value dispositions,' and a concentration/ex→11 Aug 2026How to Ask the AI: A User Perspective Survey for Large Language Model PromptingarXiv:2608.07494v1 Announce Type: cross Abstract: AI tools like ChatGPT and DeepSeek, powered by Large Language Models (LLMs), allow users to obtain instant and effective content responses simply by t→11 Aug 2026Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent HarnessesarXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harne
CompanyNVIDIA8 recent entries10 Jul 2026Jet-Long: Efficient Long-Context Extension with Dynamic Bifocal RoPEarXiv:2607.07740v1 Announce Type: cross Abstract: Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workfl→23 Jul 2026Isaac Sim-to-Real: Reinforcement Learning based Locomotion for QuadrupedsarXiv:2607.18135v1 Announce Type: cross Abstract: Learning-based approaches to locomotion have risen in popularity in recent years, showing the capability for complex legged locomotion and whole-body →24 Jul 2026NVIDIA-labs OO Agents: Native Python Object-Oriented AgentsarXiv:2607.20709v1 Announce Type: new Abstract: Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agen→24 Jul 2026Domyn-Small: A European 10B Reasoning Language ModelarXiv:2607.20448v1 Announce Type: new Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an i→28 Jul 2026Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability TestsarXiv:2607.22864v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) excel at visual interpretation but fail on spatial reasoning tasks that humans solve reliably. Existing bench→29 Jul 2026Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA KernelsarXiv:2607.24762v1 Announce Type: new Abstract: Machine learning models are increasingly embedded in everyday software, and most of their runtime is spent in a small set of compute kernels such as mat→4 Aug 2026Open-DiffLoco: Open-Source Differentiable Learning for Deployable Blind Quadruped LocomotionarXiv:2608.02069v1 Announce Type: cross Abstract: Developing deployable locomotion policies through conventional reinforcement learning often requires complex reward engineering and expensive training→11 Aug 2026EsaacSim: A Multimodal Event Camera Add-on for NVIDIA Isaac SimarXiv:2608.08522v1 Announce Type: new Abstract: Event-based vision is becoming an increasingly important sensing paradigm for robotics, yet its adoption remains limited by sensor availability and the