Tracing Agentic Failure from the Flow of Success
arXiv:2607.12747v1 Announce Type: new Abstract: Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugg
Knowledge catalogue
arXiv:2607.12747v1 Announce Type: new Abstract: Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugg
arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace
arXiv:2607.12180v1 Announce Type: cross Abstract: An AI teammate's design properties (personality, communication style, when it speaks) can shape a team's trust, coordination, and decisions. Studying
Training against GPT‑Red makes GPT‑5.6 substantially more resilient. To measure this, we replayed some of GPT‑Red’s strongest attacks—none of which our models had seen during training. GPT‑5.6 Sol pro
arXiv:2607.10744v2 Announce Type: replace Abstract: Benefiting from the powerful priors embedded in large-scale pre-training data and the emerging commonsense reasoning ability, large language models
arXiv:2607.11933v1 Announce Type: new Abstract: Cross-encoders achieve high reranking accuracy in Retrieval-Augmented Generation (RAG) pipelines but impose quadratic inference costs that limit real-ti
arXiv:2607.12612v1 Announce Type: new Abstract: BERT models have revolutionised Natural Language Processing (NLP) through their ability to process unstructured text across diverse domains. However, de
Trending AND fastest-growing in the same month? Benchmarks change daily, but only @tryramp has the database of real receipts to track this. We can confirm: demand for open-weight inference and trainin
arXiv:2607.12571v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models are deployed through pipelines that end users cannot audit, and a poisoned VLA can behave normally on clean observat
arXiv:2607.11939v1 Announce Type: new Abstract: Accurate pedestrian trajectory prediction in crowded environments remains challenging due to the multimodal uncertainty of human motion and the variable
Tune in today at 10am PST! Livestream Alert: Run ComfyUI From Claude/Cursor with Comfy MCP Host: @PurzBeats Comfy MCP lets Claude, Cursor, Amp and almost any AI agent you're already using build, run,
arXiv:2607.12372v1 Announce Type: new Abstract: Multimodal semantic segmentation (MSS) is essential for robust perception in complex environments, yet its potential remains largely untapped because of
arXiv:2607.12212v1 Announce Type: cross Abstract: Measuring retinal fluid from optical coherence tomography (OCT) drives treatment decisions in macular disease, but manual annotation is slow and segme
arXiv:2603.04113v2 Announce Type: replace-cross Abstract: Demographic attributes can be predicted from medical images, raising concerns about bias in clinical AI systems. In X-ray imaging, acquisition
arXiv:2607.12255v1 Announce Type: new Abstract: We study interaction-aware mixture-of-experts for post-stroke rigidity prediction using multi-level views of structured health records. Despite minimal
arXiv:2607.12896v1 Announce Type: new Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmen
arXiv:2607.12800v1 Announce Type: new Abstract: Learning broad world knowledge directly from raw visual data is a fundamental capability of intelligence. We introduce UniVR, the first investigation in
arXiv:2607.12861v1 Announce Type: cross Abstract: Multi-agent Reinforcement Learning (MARL) holds great potential for robot swarms, but the black-box nature of neural policies complicates strategic an
arXiv:2607.12892v1 Announce Type: cross Abstract: Modern robot learning systems increasingly rely on dense progress or value signals to evaluate intermediate states, guide policy learning, and detect
arXiv:2607.12545v1 Announce Type: cross Abstract: Adversarial robustness research has produced hundreds of defended models over the past decade, yet the literature almost universally reports robustnes
arXiv:2510.11917v2 Announce Type: replace Abstract: Dementia disorders such as Alzheimer's disease (AD) and frontotemporal dementia (FTD) exhibit overlapping electrophysiological signatures in EEG tha
arXiv:2607.12856v1 Announce Type: new Abstract: Buildings are expected to shift cooling loads in response to grid conditions. Thermal energy storage (TES) enables this shift, but scheduling it well re
arXiv:2607.12588v1 Announce Type: new Abstract: According to the recent European legislation, high-risk AI systems will have to adapt in order to comply with requirements related to specific areas, li
arXiv:2607.12959v1 Announce Type: new Abstract: LiDAR-based collaborative 3D perception in Vehicle-to-Everything (V2X) systems typically relies on fusing bird's-eye-view (BEV) features across agents.
arXiv:2607.12946v1 Announce Type: cross Abstract: Recommender-system research for Vietnamese remains limited by the absence of a public, well-documented hotel interaction resource. Building such a res
arXiv:2607.12416v1 Announce Type: new Abstract: Chromoendoscopy (CE) is a common clinical practice that sprays indigo carmine blue dye onto the gastric surface to improve the visibility of gastric les
arXiv:2606.13460v2 Announce Type: replace Abstract: Semantic 3D occupancy provides a voxelized world state for autonomous driving and robot decision making, but object and rare-class errors can affect
arXiv:2607.12756v1 Announce Type: new Abstract: Vision-language models (VLMs) process large numbers of visual tokens, resulting in substantial inference latency and memory overhead. This has motivated
arXiv:2607.12702v1 Announce Type: new Abstract: Recent advances in humanoid robotics have highlighted the importance of deployable loco-manipulation skills. Dribbling a soccer ball while evading activ
arXiv:2607.12356v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have emerged as a powerful end-to-end paradigm for robotic manipulation by mapping language instructions and 2D visu
arXiv:2607.12815v1 Announce Type: new Abstract: Chain-of-Thought (CoT) prompting is widely used as a test-time scaling strategy for Vision-Language Models (VLMs), but it remains unclear what is extend
arXiv:2512.15748v2 Announce Type: replace-cross Abstract: Visual Species Recognition (VSR) is a fundamental task in scientific disciplines that require species-level identification, including ecology,
arXiv:2607.11985v1 Announce Type: cross Abstract: Hybrid quantum-classical machine learning workflows repeatedly evaluate many small parametrized circuits during training and model exploration. In thi
arXiv:2607.12592v1 Announce Type: new Abstract: We present WanToFight, a generative game engine that simulates real-time, two-player The King of Fighters '97 (KOF~'97) gameplay from keyboard input. Pr
arXiv:2607.13003v1 Announce Type: cross Abstract: A watermark in a generative model's output is usually asked only whether a text is machine-made. The same mark can do more: attribute it to the user w
arXiv:2607.12169v1 Announce Type: new Abstract: Intercomprehension refers to partial intelligibility of an unfamiliar language (L2) by a speaker of a related language (L1). How is this zero-shot cross
We made Together GPU Clusters more reliable and easier to operate. Passive health checks, guided node repair, a rebuilt Slurm stack, better cluster visibility, external OIDC, startup scripts, and acce
We scaled a robot model natively to 8,000 timesteps of context, 5 minutes worth of muscle memory, with constant inference cost. Robot policies used to live their lives a few frames at a time (< 0.1 se
arXiv:2607.12748v1 Announce Type: cross Abstract: Farm site discovery from satellite imagery is a spatiotemporal candidate ranking problem because farm evidence is distributed across pasture, field bo
We're part of the Amazon Web Services (AWS) AI Builder Lab in New York on Friday, July 24 - a Clash of Agents competition with OpenAI, LangChain, HiddenLayer, Protopia AI, Fiddler AI, and Coder. One d
arXiv:2607.12304v1 Announce Type: new Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.
arXiv:2607.12501v1 Announce Type: new Abstract: The Forward-Forward (FF) algorithm trains each layer locally, so that a scalar goodness - the sum of squared activations - is high on real inputs and lo
arXiv:2607.12735v1 Announce Type: new Abstract: Companion work showed the grokking delay is causally the time to form task-structured representations, injectable via a contrastive prior. Here we chara
arXiv:2510.20963v2 Announce Type: replace Abstract: Multi-agent debate (MAD) was proposed as a promising approach for ensembling the wisdom of multiple large language models (LLMs) to improve reasonin
arXiv:2607.12780v1 Announce Type: cross Abstract: Quantum circuit optimization for fault-tolerant computing requires exact functional equivalence while minimizing expensive non-Clifford resources such
arXiv:2607.12248v1 Announce Type: cross Abstract: Large pretrained time-series models such as TimesFM are attractive for financial forecasting, but raw directional accuracy is a misleading scoreboard
arXiv:2607.11953v1 Announce Type: new Abstract: Does a reinforcement-learning agent that earns high reward represent its task's latent state, or only a reward-correlated shortcut? The question is usua
arXiv:2607.12790v1 Announce Type: new Abstract: Self-evolving agent systems improve by creating, revising, and retiring their own skills, but every such loop rests on a hidden assumption: a reliable e
Who touches every token that flows through @anthropic? It’s not any one model, but it’s @katelyn_lesse, @angjiang and the platform team. A year ago, it was just a messages API. Today, their platform s
arXiv:2607.12441v1 Announce Type: new Abstract: Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scientific k
arXiv:2607.12986v1 Announce Type: new Abstract: Plan evaluators can reward a strategic plan for becoming less explicit. This paper studies that failure in a staged expected-value scorer for LLM-genera
arXiv:2607.12993v1 Announce Type: new Abstract: We present X-lens, a compact feed-forward model for metric depth estimation from a variable number of calibrated fisheye and pinhole views. To support r
xai-org/grok-build, now open source xAI's grok CLI tool faced severe community backlash yesterday when it became apparent that running the command in a directory could upload that entire directory to
arXiv:2602.16918v2 Announce Type: replace-cross Abstract: We present Xray-Visual, a unified vision model architecture for large-scale image and video understanding trained on industry-scale social med
You can use OpenCode Desktop with Ollama! Try it with the top open models! Introducing Tabs OpenCode Desktop is now built around tabs. Start a new session in a tab, or open an existing session from an
OpenAI announced the launch of limited‑edition merchandise inspired by its research and deployment work, available for purchase until sold out via https://openai.com/supply/. The tweet highlighted “yo
Your coding agent doesn't need to leave the terminal to use Pinecone. We now ship official plugins and skills for the agentic IDEs and CLIs you are already building in: Claude Code, Cursor, GitHub Cop
1) If you haven't read AI as Normal Technology, these annotated slides are probably the easiest way to get a high-level overview. https://www.cs.princeton.edu/~arvindn/talks/icml-2026-annotated-slides
On July 14, 2026, OpenAI CEO Sam Altman announced a 2.5‑fold increase in usage of the company’s agentic products—Codex and ChatGPT‑based tools—within the preceding week. The tweet, which received 703.
🆕 5 Trends That Defined AI Engineering at World’s Fair 2026 https://latent.space/p/aiewf26trends @ricmac's big recap of @aidotengineer: 1. The focus shifts from agents to systems 2. Loop engineering i