AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
2 Jun 2026

Accelerating physics-informed neural networks for full waveform inversion using a hybrid quantum-classical finite-basis architecture

Model ReleasesDGX agent

arXiv:2606.01110v1 Announce Type: cross Abstract: Full waveform inversion (FWI) reconstructs heterogeneous material properties from receiver data but remains computationally demanding. Physics-informe

Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

Model ReleasesDGX agent

arXiv:2606.00920v1 Announce Type: cross Abstract: Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We

ACON: Optimizing Context Compression for Long-horizon LLM Agents


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2510.00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise re

Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.02459v1 Announce Type: new Abstract: Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is d

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

Model ReleasesDGX agent

arXiv:2512.00986v3 Announce Type: replace Abstract: A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, exi

Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation

Model ReleasesDGX agent

arXiv:2512.16310v3 Announce Type: replace-cross Abstract: LLM-based agents increasingly use multiple external tools to complete complex tasks. We study Tools Orchestration Privacy Risk (TOP-R): an age

AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

Model ReleasesDGX agent

arXiv:2606.02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future e

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

Model ReleasesDGX agent

arXiv:2603.19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large

AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

Model ReleasesDGX agent

arXiv:2606.02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical f

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

Model ReleasesDGX agent

arXiv:2603.14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reason

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

Model ReleasesDGX agent

arXiv:2606.02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, S

[AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark

Model ReleasesDGX agent

NVIDIA announced three new offerings: Cosmos 3, an advanced video generation model; Nemotron 3 Ultra, an upgraded language model; and RTX Spark, likely a tool or framework for developers. These releas

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

Model ReleasesDGX agent

arXiv:2606.00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temp

Announcing Spanner Graph algorithms: Google-grade intelligence for connected data

Model ReleasesDGX agent

At Google Cloud Next, we announced the preview of graph algorithms with Spanner Graph, bringing Google Research’s state-of-the-art graph mining capabilities natively to your database. These graph inte

Anthropic expands Project Glasswing cybersecurity program to 150 more organizations

Model ReleasesDGX agent

Anthropic PBC is expanding a program that enables organizations to test their cybersecurity defenses using its Claude Mythos Preview model. The initiative, which is known as Project Glasswing, launche

APE: Agentic Prompt Enhancer for Image Generation and Editing

Model ReleasesDGX agent

arXiv:2606.00204v1 Announce Type: new Abstract: Natural language has become a powerful interface for image generation and editing, yet text-guided visual systems remain highly sensitive to prompt form

APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL

Model ReleasesDGX agent

arXiv:2602.16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The

Approximating f-Divergences with Rank Statistics

Model ReleasesDGX agent

arXiv:2601.22784v2 Announce Type: replace-cross Abstract: We introduce a rank-statistic approximation of f-divergences that avoids explicit density-ratio estimation by working directly with the distri

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate

Model ReleasesDGX agent

arXiv:2606.00257v1 Announce Type: cross Abstract: Token-level credit assignment for language-model reinforcement learning is usually formulated as if the policy were fully trainable, while practical L

Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

Model ReleasesDGX agent

arXiv:2602.18008v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on s

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training

Model ReleasesDGX agent

arXiv:2606.00602v1 Announce Type: new Abstract: Learning transferable and interpretable representations from medical volumetric scans remains challenging due to complex anatomical structures and weak,

ASE-26: a curriculum for agentic software engineering as a discipline

Model ReleasesDGX agent

arXiv:2606.01152v1 Announce Type: cross Abstract: The work of a professional software engineer has begun to consist, increasingly, of directing agents rather than writing code, and the empirical evide

Assessment of Generative Named Entity Recognition in the Era of Large Language Models

Model ReleasesDGX agent

arXiv:2601.17898v2 Announce Type: replace Abstract: Named entity recognition (NER) is evolving from a sequence labeling task into a generative paradigm with the rise of large language models (LLMs). W

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

Model ReleasesDGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

Model ReleasesDGX agent

arXiv:2606.02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural

Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation

Model ReleasesDGX agent

arXiv:2606.02528v1 Announce Type: cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. W

AutoEval Done Right: Using Synthetic Data for Model Evaluation

Model ReleasesDGX agent

arXiv:2403.07008v3 Announce Type: replace-cross Abstract: The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data c

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

Model ReleasesDGX agent

arXiv:2606.01961v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form c

Autopilot-Preserving Residual Q-Learning with HJB-Inspired Finite-Action Risk Filtering for Fixed-Wing UAV Command Supervision

Model ReleasesDGX agent

arXiv:2606.01397v1 Announce Type: cross Abstract: A fixed-wing UAV must hold airspeed, altitude, and heading references under wind, gusts, and turbulence, channels coupled so that correcting one can d

AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2606.00671v1 Announce Type: new Abstract: We present AXIOM, a trust-first neuro-symbolic execution architecture for natural-language mathematical reasoning. In AXIOM, the language model function

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

Model ReleasesDGX agent

arXiv:2606.02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approac

Bayesian Inference of Nonlinear Malaria Dynamics in Ghana via an Ensemble Markov Chain Monte Carlo Sampler

Model ReleasesDGX agent

arXiv:2606.00783v1 Announce Type: cross Abstract: Reliable quantification of malaria dynamics in sub-Saharan Africa is hindered by short, noisy, and spatially heterogeneous surveillance records. In Gh

Before and After Temperature: A Distributional View of Creative LLM Generation

Model ReleasesDGX agent

arXiv:2606.01451v1 Announce Type: new Abstract: Reference-free evaluation of large language model (LLM) creativity relies on perplexity, entropy, and top-1 margin. We show that a much stronger signal

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

Model ReleasesDGX agent

arXiv:2606.01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differen

Benchmark Dataset for Catalysis on 2D MXenes

Model ReleasesDGX agent

arXiv:2606.00794v1 Announce Type: cross Abstract: Merging first-principles calculations with machine learning (ML), we aim to accelerate the exploration of catalytic behaviour in novel materials. We f

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

Model ReleasesDGX agent

arXiv:2505.24621v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarkin

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

Model ReleasesDGX agent

arXiv:2606.01629v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. L

Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade Hardware

Model ReleasesDGX agent

arXiv:2606.01338v1 Announce Type: new Abstract: Biopharmaceutical manufacturing organizations operate under regulatory frameworks such as FDA guidance, EU Good Manufacturing Practice (GMP), and the EU

Benchmarking Multimodal LLMs on Code Generation for Complex Interactive Webpages

Model ReleasesDGX agent

arXiv:2606.00154v1 Announce Type: cross Abstract: Recent advancements in multimodal large language models (MLLMs) have achieved remarkable progress in multimodal reasoning and code generation, catalyz

Benchmarking Recursive-Collapse Warning Claims Under Matched False-Positive Control

Model ReleasesDGX agent

arXiv:2606.00329v1 Announce Type: cross Abstract: Recursive systems can enter collapse-like regimes -- self-reinforcing amplification, persistent recursion, and narrowing diversity that mask accelerat

Benchmarking Security Risk Detection and Verification in Open Agentic Skill Ecosystems

Model ReleasesDGX agent

arXiv:2606.00925v1 Announce Type: cross Abstract: Open agent platforms allow community contributors to publish reusable skills that agents can invoke at runtime. This extensibility also creates a supp

Benchmarking Waitlist Mortality Prediction in Heart Transplantation Through Time-to-Event Modeling using New Longitudinal UNOS Dataset

Model ReleasesDGX agent

arXiv:2507.07339v2 Announce Type: replace-cross Abstract: Decisions about managing patients on the heart transplant waitlist are currently made by committees of doctors who consider multiple factors,

Benchmarks for Vision-Language Models in Urban Perception Should Be Reliability-Aware and Negotiated

Model ReleasesDGX agent

arXiv:2606.00871v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to generate structured descriptions of street-level imagery for tasks such as streetscape auditing

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes

Model ReleasesDGX agent

arXiv:2606.02215v1 Announce Type: new Abstract: Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social p

Beware of the Batch Size: Hyperparameter Bias in Evaluating LoRA

Model ReleasesDGX agent

arXiv:2602.09492v2 Announce Type: replace-cross Abstract: Low-rank adaptation (LoRA) is a standard approach for fine-tuning large language models, yet its many variants report conflicting empirical ga

Beyond ell_2-norm and ell_infty-norm: A Curvature-Inspired ell_p-Norm Scheme for Deep Neural Networks

Model ReleasesDGX agent

arXiv:2606.02078v1 Announce Type: new Abstract: The existing optimizers for deep neural networks (DNNs) typically rely on either the ell_2 norm or the ell_infty norm, resulting in optimizers that do n

Beyond Isolated Behaviors: Hierarchical User Modeling for LLM Personalization

Model ReleasesDGX agent

arXiv:2606.02300v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, yet personalizing their outputs to individual users remai

Beyond Rigid: Benchmarking Non-Rigid Video Editing

Model ReleasesDGX agent

arXiv:2601.18340v2 Announce Type: replace Abstract: As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance f

Beyond Scalar Rewards: Dense Feedback for LLM Policy Synthesis in Sequential Social Dilemmas

Model ReleasesDGX agent

arXiv:2603.19453v2 Announce Type: replace Abstract: We study LLM policy synthesis: using a language model to iteratively generate programmatic agent policies for multi-agent environments. Rather than

Beyond Semantic Understanding: Preserving Collaborative Frequency Components in LLM-based Recommendation

Model ReleasesDGX agent

arXiv:2508.10312v2 Announce Type: replace Abstract: Recommender systems in concert with Large Language Models (LLMs) present promising avenues for generating semantically-informed recommendations. How

Beyond Static Gaussians: An Empirical Investigation of Architectural Paradigms for Dynamic 3D Scene Reconstruction

Model ReleasesDGX agent

arXiv:2606.00452v1 Announce Type: new Abstract: Dynamic scene reconstruction via 3D Gaussian Splatting (3DGS) has emerged as a compelling approach for representing evolving environments, yet understan

Beyond Text and Tables: Vision-Language Model Integration in ComProScanner for Extracting Materials Data from Scientific Figures with High Accuracy

Model ReleasesDGX agent

arXiv:2606.00065v1 Announce Type: cross Abstract: Automated extraction of materials composition-property data from scientific literature has advanced considerably with the development of large languag

Beyond the Simplex: Balanced Prototype Geometry for Scorer-Agnostic Open-Set Recognition

Model ReleasesDGX agent

arXiv:2606.01883v1 Announce Type: cross Abstract: Open-set recognition (OSR) requires a classifier to reject inputs from unseen classes which is essential in safety-critical settings such as medical i

Big paper on AI coding agents using Github & other data The auto-complete tools (Copilot) led to 2.2x more code, local agents like original …

Model ReleasesDGX agent

Big paper on AI coding agents using Github & other data The auto-complete tools (Copilot) led to 2.2x more code, local agents like original Claude Code led to 7.4x, & current remote coding agents 17.3

BLISS: A Lightweight Bilevel Influence Scoring Method for Data Selection in Language Model Pretraining

Model ReleasesDGX agent

arXiv:2510.06048v4 Announce Type: replace Abstract: Effective data selection is essential for pretraining large language models (LLMs), enhancing efficiency and improving generalization to downstream

Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention

Model ReleasesDGX agent

arXiv:2512.10414v2 Announce Type: replace Abstract: Recently, reinforcement learning (RL) has become a common choice in enhancing the reasoning capabilities of vision-language models (VLMs). Consideri

Boundary-Protection W8A8 HiFloat8 Quantization for Large-Scale Text-to-Video Diffusion Transformers

Model ReleasesDGX agent

arXiv:2606.00957v1 Announce Type: new Abstract: We present a post-training quantization (PTQ) approach for Wan2.1-T2V-14B, a 14-billion-parameter text-to-video diffusion transformer, targeting the W8A

BraveGuard: From Open-World Threats to Safer Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.01166v1 Announce Type: cross Abstract: Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shi

Bridging Requirements and Architecture: Multi-Agent Orchestration with External Knowledge and Hierarchical Memory

Model ReleasesDGX agent

arXiv:2606.01385v1 Announce Type: cross Abstract: Software architecture design is a critical yet inherently complex and knowledge-intensive phase that requires balancing competing quality attributes a

Bridging the Sim-to-Real Gap in Semiconductor Visual Program Synthesis via Input Binarization

Model ReleasesDGX agent

arXiv:2606.02434v1 Announce Type: new Abstract: Precise parametric control over circuit geometry is essential for semiconductor inspection, yet obtaining sufficient real training data remains costly.

← Previous
1…178179180181182…377
Next →