General Agent Evaluation
arXiv:2602.22953v2 Announce Type: replace Abstract: General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measur
Knowledge catalogue
arXiv:2602.22953v2 Announce Type: replace Abstract: General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measur
arXiv:2605.08838v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) is widely used to augment large language models (LLMs) with external knowledge. However, many benchmark datasets,
arXiv:2605.10115v1 Announce Type: new Abstract: Tackling the task of materials generation, we aim to enhance the previously proposed All-atom Diffusion Transformer (ADiT) by introducing SymADiT, a sym
arXiv:2605.08733v1 Announce Type: new Abstract: Expressive generative policies such as diffusion and flow models are appealing for MaxEnt online reinforcement learning because of their ability to mode
arXiv:2605.10291v1 Announce Type: cross Abstract: Recent advances in generative artificial intelligence (AI) are reshaping who enters entrepreneurship, but not who reaches the top of the quality distr
arXiv:2605.10645v1 Announce Type: new Abstract: Data-driven medical AI is traditionally formulated as a discriminative mapping from input X to output Y via a learned function f, which does not general
arXiv:2605.08851v1 Announce Type: cross Abstract: The scarcity of high-quality imaging data for coronary angiography (CAG) stenosis limits the clinical translation of automated stenosis detection. Syn
arXiv:2605.08392v1 Announce Type: new Abstract: Practical diffusion sampling is a numerical approximation problem: under a fixed inference budget, one must simulate a reverse-time ODE or SDE using onl
arXiv:2605.09608v1 Announce Type: new Abstract: Continual post-training aims to extend large language models (LLMs) with new knowledge, skills, and behaviors, yet it remains unclear when sequential up
arXiv:2605.08109v1 Announce Type: new Abstract: Inertial microfluidic devices (IMDs) offer low-cost, high-throughput alternative techniques for many traditional particle- (or cell-) manipulation tasks
arXiv:2605.10739v1 Announce Type: cross Abstract: We introduce SMART-HC-VQA, a Sentinel-2-based visual question answering dataset derived from the IARPA SMART Heavy Construction dataset, designed for
arXiv:2509.20863v3 Announce Type: replace Abstract: Diffusion models have recently shown strong potential in language modeling, offering faster generation compared to traditional autoregressive approa
arXiv:2605.10108v1 Announce Type: new Abstract: Joint named entity recognition (NER) and relation extraction (RE) is a fundamental task in natural language processing for constructing knowledge graphs
arXiv:2605.09973v1 Announce Type: cross Abstract: Reliable detection of personally identifiable information (PII) is increasingly important across modern data-processing systems, yet the task remains
arXiv:2509.17815v2 Announce Type: replace Abstract: Global optimization, particularly for non-convex functions with multiple local minima, poses significant challenges for traditional gradient-based m
arXiv:2603.12275v1 Announce Type: cross Abstract: Unlearning knowledge is a pressing and challenging task in Large Language Models (LLMs) because of their unprecedented capability to memorize and dige
Allison Johnson / The Verge: Google announces Gemini Intelligence, bundling existing and new Gemini features, including task automation across apps and vibe-coding own Android widgets — We're one ste
Google LLC is upgrading its consumer device portfolio with a set of Android features called Gemini Intelligence and a new laptop series. The company debuted the products today during a virtual event c
Google DeepMind: Google DeepMind details a Gemini-powered mouse pointer that understands what it is pointing at, allowing users to perform tasks without using text-heavy prompts — We are developing mo
Tim Starks / CyberScoop: Google launches Intrusion Logging, an Android feature developed in partnership with Amnesty International and others, on Android 16 Pixel devices for now — Intrusion Logging m
arXiv:2512.04475v5 Announce Type: replace-cross Abstract: Machine learning on graphs has made substantial progress across domains such as molecular property prediction and chip design. Yet benchmarkin
arXiv:2605.09408v1 Announce Type: new Abstract: Link prediction (inferring missing or future connections between nodes in a graph) is a fundamental problem in network science with widespread applicati
arXiv:2605.10893v1 Announce Type: new Abstract: Large vision-language models suffer from visual ungroundedness: they can produce a fluent, confident, and even correct response driven entirely by langu
arXiv:2603.16253v2 Announce Type: replace-cross Abstract: Vision-language process reward models (VL-PRMs) are increasingly used to score intermediate reasoning steps and rerank candidates under test-t
arXiv:2508.20325v3 Announce Type: replace-cross Abstract: As Large Language Models (LLMs) become increasingly integral to various domains, their potential to generate harmful responses has prompted si
arXiv:2411.04077v2 Announce Type: replace Abstract: By leveraging both texts and images, large vision language models (LVLMs) have shown significant progress in various multi-modal tasks. Nevertheless
arXiv:2605.09502v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting assumes that generated reasoning reflects a model's internal computation. We show this assumption is wrong in a speci
arXiv:2605.09972v1 Announce Type: cross Abstract: End-to-end autonomous driving has witnessed rapid progress, yet existing benchmarks are increasingly saturated, with state-of-the-art models achieving
arXiv:2410.14927v2 Announce Type: replace-cross Abstract: Automated equity trading requires converting noisy market and news signals into executable portfolio decisions under risk, turnover, and trans
arXiv:2605.09465v1 Announce Type: new Abstract: High-precision heavy-duty grading is a common step in earthworks, traditionally carried out manually by skilled operators. Removing a significant amount
arXiv:2605.08864v1 Announce Type: new Abstract: We study online estimation in latent-variable models by recasting the problem as tracking a moving empirical equilibrium. Standard online EM and stochas
arXiv:2605.10546v1 Announce Type: new Abstract: Pixel-based deep reinforcement learning agents are typically trained on heavily downsampled visual observations, a convention inherited from early bench
arXiv:2605.08665v1 Announce Type: new Abstract: Large reasoning models achieve high accuracy through extended chain-of-thought but generate 5--8 more tokens than necessary, applying verbose reasoning
arXiv:2404.18923v5 Announce Type: replace Abstract: We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic
arXiv:2605.09348v1 Announce Type: cross Abstract: Large Language Models (LLMs) provide flexible natural language processing capabilities, while knowledge graphs (KGs) offer explicit and structured kno
arXiv:2605.08143v1 Announce Type: cross Abstract: Large language models encode vast factual knowledge that inevitably becomes outdated or incorrect after deployment, yet retraining is costly prohibiti
NVIDIA engineers and researchers utilize OpenAI's Codex, a large language model trained on code, to accelerate software development and improve productivity across their engineering workflows. The art
arXiv:2605.09523v1 Announce Type: new Abstract: Neural operators provide fast surrogate models for time-dependent partial differential equations, but their standard autoregressive use usually assumes
arXiv:2605.08538v1 Announce Type: new Abstract: Current LLM agents lack principled mechanisms for managing persistent memory across long interaction horizons. We present a biologically-grounded memory
arXiv:2605.08533v1 Announce Type: new Abstract: Clinical decision-making in emergency medicine demands rapid, accurate diagnoses under uncertainty. Despite benchmark progress, evidence for LLMs as int
arXiv:2501.12202v4 Announce Type: replace Abstract: We present Hunyuan3D 2.0, an advanced large-scale 3D synthesis system for generating high-resolution textured 3D assets. This system includes two fo
arXiv:2604.15113v2 Announce Type: replace Abstract: Vector Symbolic Architectures (VSAs) provide a well-defined algebraic framework for compositional representations in hyperdimensional spaces. We int
arXiv:2605.10009v1 Announce Type: new Abstract: Query-based image retrieval (QBIR) requires retrieving relevant images given diverse and often stylistically heterogeneous queries, such as sketches, ar
I love seeing a new eval with such low scores. When we announced GPT-5.5, almost every benchmark had a score above 50%. It's time to retire evals like GQPA and bring in a new set. The first ProgramBen
I needed to book flights for a bunch of upcoming travel. As always, I used Claude Cowork to do it. In the past, Cowork has been decent at booking flights, but with Opus 4.7, for the first time ever, i
I put my flight preferences in my Cowork instructions, then let Opus get to work. It opened my browser, navigated a bunch of websites, and booked everything for me. The result: Cowork booked 8 flights
arXiv:2403.18136v3 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) have gained popularity in numerous domains, yet they are vulnerable to backdoor attacks that can compromise their
arXiv:2602.22551v2 Announce Type: replace-cross Abstract: Cancer is often driven by specific combinations of an estimated two to nine gene mutations, known as multi-hit combinations. Identifying these
Palo Alto Networks Inc. today launched Idira, a new identity security platform designed to manage human, machine and artificial intelligence agent identities across the enterprise under a single privi
arXiv:2605.08841v1 Announce Type: new Abstract: Vision-Language Models (VLMs) exhibit systematic bias toward visual illusions, recalling memorized facts rather than perceiving actual visual difference
arXiv:2605.09256v1 Announce Type: cross Abstract: We introduce a use of the (M)-cover (or (M)-layer) transform for machine learning. The method replicates a model (M) times, but instead of coupling th
arXiv:2605.08184v1 Announce Type: cross Abstract: This research addresses a validated TMS EEG cleaning pipeline and a corresponding benchmark dataset. It evaluates two widely used artifact removal pip
'In California, every professor has 10 different startups with their students. In some Canadian universities, it's one professor out of ten.' @jpineau1, Chief AI Officer at @Cohere at #WebSummitVancou
arXiv:2605.08295v1 Announce Type: cross Abstract: While random demonstration labels barely hurt in-context learning (Min et al., 2022), we show that homogeneous labels--even semantically valid ones--c
arXiv:2601.16097v2 Announce Type: replace Abstract: Large Language Models enable users to access database using natural language interfaces using tools like Text2SQL, Text2SPARQL, and Text2Cypher, whi
arXiv:2605.10267v1 Announce Type: new Abstract: In industrial procurement, an LLM answer is useful only if it survives a standards check: recommended material must match operating condition, every par
arXiv:2605.10039v1 Announce Type: cross Abstract: Frontier coding agents read configuration files (CLAUDE.md, AGENTS.md, Cursor Rules) at session start and are expected to follow the conventions insid
arXiv:2605.09870v1 Announce Type: cross Abstract: We propose SVAR-FM (Structural VAR with Flow Matching), a framework for time series causal discovery that treats a physics-based simulator as a mechan
arXiv:2605.08233v1 Announce Type: cross Abstract: Inverse design of RF passive components from S-parameters is a high-dimensional, ill-posed problem, and prior generative approaches are limited to sin
arXiv:2605.08664v1 Announce Type: new Abstract: Current image quality assessment methods are heavily biased towards global distortions (e.g., noise, blur), neglecting local perceptual artifacts such a