AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,805 results
Model Releases

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-…

DGX agent

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger int

model-releasesopenai--x
3 Jun 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

DGX agent

arXiv:2606.02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceed

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

What Makes Interaction Trajectories Effective for Training Terminal Agents?

DGX agent

arXiv:2606.03461v1 Announce Type: new Abstract: Stronger code agents are commonly assumed to be superior teachers for post-training, yet this assumption remains poorly disentangled from task difficult

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

When Model Merging Breaks Routing: Training-Free Calibration for MoE

DGX agent

arXiv:2606.03391v1 Announce Type: cross Abstract: Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing mergi

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation

DGX agent

arXiv:2606.03532v1 Announce Type: cross Abstract: Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- whi

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

DGX agent

arXiv:2606.03837v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it a

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

DGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

DGX agent

arXiv:2602.08873v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isola

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Will Accurate Fields Mislead Photonic Design? FromGlobal Accuracy to Port Readout

DGX agent

arXiv:2606.03038v1 Announce Type: new Abstract: Neural field surrogates can accelerate photonic design loops, but a surrogate that looks accurate in global field error can still mis-rank candidate dev

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

DGX agent

arXiv:2503.07265v4 Announce Type: replace-cross Abstract: Text-to-Image (T2I) models are capable of generating high-quality artistic creations and visual content. However, existing research and evalua

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Wordle 1,809 5/6 🟨⬛🟨⬛⬛ ⬛🟨⬛⬛⬛ ⬛🟩🟨🟨⬛ 🟩🟩🟩🟩⬛ 🟩🟩🟩🟩🟩

DGX agent

This post documents a completed Wordle puzzle (puzzle #1,809) solved in five attempts, showing the progression of letter guesses through color-coded feedback (yellow for correct letters in wrong posit

model-releasesanthropic--x
3 Jun 2026
Model Releases

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

DGX agent

arXiv:2606.02908v1 Announce Type: cross Abstract: Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute val

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

3D Segment Anything Model with Visual Mamba for Diagnosing Placenta Accreta Spectrum

DGX agent

arXiv:2606.00489v1 Announce Type: new Abstract: Placenta Accreta Spectrum (PAS) is a rare but highly dangerous obstetric disease. Early and accurate PAS diagnosis is critical for maternal health. Trad

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

3DCodeBench: Benchmarking Agentic Procedural 3D Modeling Via Code

DGX agent

arXiv:2606.01057v1 Announce Type: cross Abstract: Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neur

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

3rd Place at CVPR 2026 CASTLE Challenge: Agentic Multi-View Long-Context Video Understanding via Hierarchical Knowledge Graph Retrieval

DGX agent

arXiv:2606.01933v1 Announce Type: new Abstract: This paper presents our winning methodology for the CASTLE 2026 Challenge at the CVPR 2026 EgoVis Workshop, where our team secured third place globally.

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

A Closer Look at In-Distribution vs. Out-of-Distribution Accuracy for Open-Set Test-time Adaptation

DGX agent

arXiv:2606.01973v1 Announce Type: cross Abstract: Open-set test-time adaptation (TTA) updates models on new data in the presence of input shifts and unknown output classes. While recent methods have m

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

A Comparative Analysis of Machine Learning Algorithms for Multi-Task Prediction of the Parameters of the Pectin Hydrolysis--Extraction Process

DGX agent

arXiv:2606.00821v1 Announce Type: new Abstract: This study addresses the challenge of controlling a complex, multi-parameter technological process -- pectin hydrolysis--extraction -- using machine lea

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

A Finite-Calibration Regime Map for LLM Judge Panels

DGX agent

arXiv:2606.01034v1 Announce Type: new Abstract: We study when LLM judge panels should be calibrated with low-dimensional stackers versus joint output tables under finite human-label budgets. Low-dimen

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

DGX agent

arXiv:2606.02398v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation,

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

A Machine-to-Machine Knowledge-Guided LLM Agent for Generalizable Radiotherapy Treatment Planning

DGX agent

arXiv:2606.00922v1 Announce Type: cross Abstract: In this work, we propose a prototype machine-to-machine (M2M) knowledge-guided Large Language Model (LLM) framework for automated radiotherapy treatme

model-releasesarxiv-cs-ro
2 Jun 2026
Model Releases

A Methodological Framework for Explicit Control of the Speed-Accuracy Trade-off in Brain-Computer Interfaces

DGX agent

arXiv:2606.00106v1 Announce Type: cross Abstract: Brain-computer interfaces (BCIs) are limited by low signal-to-noise ratio in modalities such as electroencephalography, which requires multiple trials

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

DGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Novel Data Augmentation Strategy for Robust Deep Learning Classification of Biomedical Time-Series Data: Application to ECG and EEG Analysis

DGX agent

arXiv:2507.12645v1 Announce Type: cross Abstract: The increasing need for accurate and unified analysis of diverse biological signals, such as ECG and EEG, is paramount for comprehensive patient asses

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Per-Component Diagnostic Protocol for Neural HJB-PIDE Solvers under Control-Dependent Levy Jumps

DGX agent

arXiv:2606.01122v1 Announce Type: new Abstract: We propose a five-step diagnostic protocol for residual-trained neural HJB-PIDE solvers with control-dependent Levy jumps, targeting a general failure m

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision

DGX agent

arXiv:2606.01992v1 Announce Type: cross Abstract: Industrial anomaly detection has historically been a unimodal task. Recent multimodal vision-language models have produced systems that admit textual

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Systematic Benchmark of Intraoperative Ultrasound-to-MR Synthesis for Brain Tumour Surgery

DGX agent

arXiv:2606.00630v1 Announce Type: new Abstract: Intraoperative ultrasound (ioUS) is a versatile, cost-effective modality in brain tumour surgery, but its interpretation is difficult: acquisition plane

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research

DGX agent

arXiv:2507.08038v3 Announce Type: replace-cross Abstract: Language model agents are increasingly used to automate scientific research, yet evaluating their scientific contributions remains a challenge

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Absorbing Complexity: An Interaction-Native Knowledge Harness for Financial LLM Agents

DGX agent

arXiv:2606.01886v1 Announce Type: new Abstract: Financial AI agents often fail for a simple reason: they make users carry the complexity. A user must repeatedly restate goals, risk preferences, portfo

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Accelerating data lakes: Optimizing Apache Iceberg and Spark with gcs-analytics-core

DGX agent

Many data engineers spend significant time managing compatibility and getting best performance across multiple analytics engines. To help solve this pain point, we are excited to announce gcs-analytic

model-releasesgoogle-cloud-ai
2 Jun 2026
Model Releases

Accelerating physics-informed neural networks for full waveform inversion using a hybrid quantum-classical finite-basis architecture

DGX agent

arXiv:2606.01110v1 Announce Type: cross Abstract: Full waveform inversion (FWI) reconstructs heterogeneous material properties from receiver data but remains computationally demanding. Physics-informe

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

DGX agent

arXiv:2606.00920v1 Announce Type: cross Abstract: Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

ACON: Optimizing Context Compression for Long-horizon LLM Agents

DGX agent

arXiv:2510.00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise re

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

DGX agent

arXiv:2606.02459v1 Announce Type: new Abstract: Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is d

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

DGX agent

arXiv:2512.00986v3 Announce Type: replace Abstract: A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, exi

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation

DGX agent

arXiv:2512.16310v3 Announce Type: replace-cross Abstract: LLM-based agents increasingly use multiple external tools to complete complex tasks. We study Tools Orchestration Privacy Risk (TOP-R): an age

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

DGX agent

arXiv:2606.02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future e

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

DGX agent

arXiv:2603.19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

DGX agent

arXiv:2606.02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical f

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

DGX agent

arXiv:2603.14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reason

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

DGX agent

arXiv:2606.02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, S

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

[AINews] NVIDIA Cosmos 3, Nemotron 3 Ultra, and RTX Spark

DGX agent

NVIDIA announced three new offerings: Cosmos 3, an advanced video generation model; Nemotron 3 Ultra, an upgraded language model; and RTX Spark, likely a tool or framework for developers. These releas

model-releaseslatent-space
2 Jun 2026
Model Releases

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

DGX agent

arXiv:2606.00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temp

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Announcing Spanner Graph algorithms: Google-grade intelligence for connected data

DGX agent

At Google Cloud Next, we announced the preview of graph algorithms with Spanner Graph, bringing Google Research’s state-of-the-art graph mining capabilities natively to your database. These graph inte

model-releasesgoogle-cloud-ai
2 Jun 2026
Model Releases

Anthropic expands Project Glasswing cybersecurity program to 150 more organizations

DGX agent

Anthropic PBC is expanding a program that enables organizations to test their cybersecurity defenses using its Claude Mythos Preview model. The initiative, which is known as Project Glasswing, launche

model-releasessiliconangle
2 Jun 2026
Model Releases

APE: Agentic Prompt Enhancer for Image Generation and Editing

DGX agent

arXiv:2606.00204v1 Announce Type: new Abstract: Natural language has become a powerful interface for image generation and editing, yet text-guided visual systems remain highly sensitive to prompt form

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL

DGX agent

arXiv:2602.16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Approximating f-Divergences with Rank Statistics

DGX agent

arXiv:2601.22784v2 Announce Type: replace-cross Abstract: We introduce a rank-statistic approximation of f-divergences that avoids explicit density-ratio estimation by working directly with the distri

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate

DGX agent

arXiv:2606.00257v1 Announce Type: cross Abstract: Token-level credit assignment for language-model reinforcement learning is usually formulated as if the policy were fully trainable, while practical L

model-releasesarxiv-cs-ai
2 Jun 2026
← Previous
1…226227228229230…476
Next →