AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

A Closer Look at In-Distribution vs. Out-of-Distribution Accuracy for Open-Set Test-time Adaptation

DGX agent

arXiv:2606.01973v1 Announce Type: cross Abstract: Open-set test-time adaptation (TTA) updates models on new data in the presence of input shifts and unknown output classes. While recent methods have m

model-releasesarxiv-cs-cv
2 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

A Comparative Analysis of Machine Learning Algorithms for Multi-Task Prediction of the Parameters of the Pectin Hydrolysis--Extraction Process

DGX agent

arXiv:2606.00821v1 Announce Type: new Abstract: This study addresses the challenge of controlling a complex, multi-parameter technological process -- pectin hydrolysis--extraction -- using machine lea

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

A Finite-Calibration Regime Map for LLM Judge Panels

DGX agent

arXiv:2606.01034v1 Announce Type: new Abstract: We study when LLM judge panels should be calibrated with low-dimensional stackers versus joint output tables under finite human-label budgets. Low-dimen

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL

DGX agent

arXiv:2606.02398v1 Announce Type: cross Abstract: Reinforcement learning (RL) post-training improves large language models (LLMs) on individual domains such as mathematical reasoning, code generation,

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

A Machine-to-Machine Knowledge-Guided LLM Agent for Generalizable Radiotherapy Treatment Planning

DGX agent

arXiv:2606.00922v1 Announce Type: cross Abstract: In this work, we propose a prototype machine-to-machine (M2M) knowledge-guided Large Language Model (LLM) framework for automated radiotherapy treatme

model-releasesarxiv-cs-ro
2 Jun 2026
Model Releases

A Methodological Framework for Explicit Control of the Speed-Accuracy Trade-off in Brain-Computer Interfaces

DGX agent

arXiv:2606.00106v1 Announce Type: cross Abstract: Brain-computer interfaces (BCIs) are limited by low signal-to-noise ratio in modalities such as electroencephalography, which requires multiple trials

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Multi-Domain Red Teaming Framework for Safety, Robustness, and Fairness Evaluation of Medical Large Language Models

DGX agent

arXiv:2606.00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Novel Data Augmentation Strategy for Robust Deep Learning Classification of Biomedical Time-Series Data: Application to ECG and EEG Analysis

DGX agent

arXiv:2507.12645v1 Announce Type: cross Abstract: The increasing need for accurate and unified analysis of diverse biological signals, such as ECG and EEG, is paramount for comprehensive patient asses

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Per-Component Diagnostic Protocol for Neural HJB-PIDE Solvers under Control-Dependent Levy Jumps

DGX agent

arXiv:2606.01122v1 Announce Type: new Abstract: We propose a five-step diagnostic protocol for residual-trained neural HJB-PIDE solvers with control-dependent Levy jumps, targeting a general failure m

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

A Structured Benchmark for Text-Guided Anomaly Detection: When Language Stops Conditioning the Decision

DGX agent

arXiv:2606.01992v1 Announce Type: cross Abstract: Industrial anomaly detection has historically been a unimodal task. Recent multimodal vision-language models have produced systems that admit textual

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

A Systematic Benchmark of Intraoperative Ultrasound-to-MR Synthesis for Brain Tumour Surgery

DGX agent

arXiv:2606.00630v1 Announce Type: new Abstract: Intraoperative ultrasound (ioUS) is a versatile, cost-effective modality in brain tumour surgery, but its interpretation is difficult: acquisition plane

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research

DGX agent

arXiv:2507.08038v3 Announce Type: replace-cross Abstract: Language model agents are increasingly used to automate scientific research, yet evaluating their scientific contributions remains a challenge

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Absorbing Complexity: An Interaction-Native Knowledge Harness for Financial LLM Agents

DGX agent

arXiv:2606.01886v1 Announce Type: new Abstract: Financial AI agents often fail for a simple reason: they make users carry the complexity. A user must repeatedly restate goals, risk preferences, portfo

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Accelerating physics-informed neural networks for full waveform inversion using a hybrid quantum-classical finite-basis architecture

DGX agent

arXiv:2606.01110v1 Announce Type: cross Abstract: Full waveform inversion (FWI) reconstructs heterogeneous material properties from receiver data but remains computationally demanding. Physics-informe

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Accuracy, Stability, and Repeated-Run Reliability of Large Language Models on Deterministic Programming Tasks

DGX agent

arXiv:2606.00920v1 Announce Type: cross Abstract: Run-level pass rate overstates retry-free coverage by up to 17.8 percentage points -- and the gap is largest precisely for mid-performing systems. We

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

ACON: Optimizing Context Compression for Long-horizon LLM Agents

DGX agent

arXiv:2510.00615v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed as agents in dynamic real-world environments, where success depends on maintaining precise re

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Active Exploring like a Pigeon: Reinforcing Spatial Reasoning via Agentic Vision-Language Models

DGX agent

arXiv:2606.02459v1 Announce Type: new Abstract: Enabling Vision-Language Models (VLMs) to perform spatial reasoning remains challenging. Existing approaches treat VLMs as passive observers, which is d

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

ADRA-Bank: A Modular Benchmark for Academic Deep Research Agents

DGX agent

arXiv:2512.00986v3 Announce Type: replace Abstract: A surge in academic publications calls for automated deep research (DR) systems, but accurately evaluating them is still an open problem. First, exi

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation

DGX agent

arXiv:2512.16310v3 Announce Type: replace-cross Abstract: LLM-based agents increasingly use multiple external tools to complete complex tasks. We study Tools Orchestration Privacy Risk (TOP-R): an age

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AGENTCL: Toward Rigorous Evaluation of Continual Learning in Language Agents

DGX agent

arXiv:2606.02461v1 Announce Type: new Abstract: Language agents spend substantial inference time solving individual tasks, yet the experience acquired in one episode is often underutilized in future e

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science

DGX agent

arXiv:2603.19005v2 Announce Type: replace-cross Abstract: Data science plays a critical role in transforming complex data into actionable insights across numerous domains. Recent developments in large

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentPLM: Agentic Protein Language Models with Reasoning-Augmented Decoding for Protein Sequence Design

DGX agent

arXiv:2606.02386v1 Announce Type: new Abstract: Protein language models (PLMs) are passive oracles: they generate sequences in a single forward pass with no mechanism to consult external biophysical f

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentProcessBench: Diagnosing Step-Level Process Quality in Tool-Using Agents

DGX agent

arXiv:2603.14465v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have evolved into tool-using agents, they remain brittle in long-horizon interactions. Unlike mathematical reason

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations

DGX agent

arXiv:2606.02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, S

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation

DGX agent

arXiv:2606.00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temp

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

APE: Agentic Prompt Enhancer for Image Generation and Editing

DGX agent

arXiv:2606.00204v1 Announce Type: new Abstract: Natural language has become a powerful interface for image generation and editing, yet text-guided visual systems remain highly sensitive to prompt form

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

APEX-SQL: Talking to the data via Agentic Exploration for Text-to-SQL

DGX agent

arXiv:2602.16720v2 Announce Type: replace-cross Abstract: Text-to-SQL systems powered by Large Language Models have excelled on academic benchmarks but struggle in complex enterprise environments. The

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Approximating f-Divergences with Rank Statistics

DGX agent

arXiv:2601.22784v2 Announce Type: replace-cross Abstract: We introduce a rank-statistic approximation of f-divergences that avoids explicit density-ratio estimation by working directly with the distri

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

ARCA: Adapter-Residual Credit Assignment When Token Signals Degenerate

DGX agent

arXiv:2606.00257v1 Announce Type: cross Abstract: Token-level credit assignment for language-model reinforcement learning is usually formulated as if the policy were fully trainable, while practical L

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework

DGX agent

arXiv:2602.18008v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise in constructing mechanistic models from data. However, existing evaluations largely focus on s

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

ASAP: Advancing Medical Volumetric Representation Learning with Anatomy-aware Semantically-adaptive Pre-training

DGX agent

arXiv:2606.00602v1 Announce Type: new Abstract: Learning transferable and interpretable representations from medical volumetric scans remains challenging due to complex anatomical structures and weak,

model-releasesarxiv-cs-cv
2 Jun 2026
Model Releases

ASE-26: a curriculum for agentic software engineering as a discipline

DGX agent

arXiv:2606.01152v1 Announce Type: cross Abstract: The work of a professional software engineer has begun to consist, increasingly, of directing agents rather than writing code, and the empirical evide

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Assessment of Generative Named Entity Recognition in the Era of Large Language Models

DGX agent

arXiv:2601.17898v2 Announce Type: replace Abstract: Named entity recognition (NER) is evolving from a sequence labeling task into a generative paradigm with the rise of large language models (LLMs). W

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

ATLAS: Agentic Test-time Learning-to-Allocate Scaling

DGX agent

arXiv:2606.01667v1 Announce Type: new Abstract: Test-time scaling has become a major way to improve large language model reasoning, but its orchestration has remained designer-engineered: a fixed samp

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

DGX agent

arXiv:2606.02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Auditing Asset-Specific Preferences in Financial Large Language Models: Evidence from Bitcoin Representations and Portfolio Allocation

DGX agent

arXiv:2606.02528v1 Announce Type: cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. W

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

AutoEval Done Right: Using Synthetic Data for Model Evaluation

DGX agent

arXiv:2403.07008v3 Announce Type: replace-cross Abstract: The evaluation of machine learning models using human-labeled validation data can be expensive and time-consuming. AI-labeled synthetic data c

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models

DGX agent

arXiv:2606.01961v1 Announce Type: new Abstract: Autonomous agents are increasingly expected to support end-to-end medical-AI research workflows, moving beyond isolated prediction tasks or short-form c

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Autopilot-Preserving Residual Q-Learning with HJB-Inspired Finite-Action Risk Filtering for Fixed-Wing UAV Command Supervision

DGX agent

arXiv:2606.01397v1 Announce Type: cross Abstract: A fixed-wing UAV must hold airspeed, altitude, and heading references under wind, gusts, and turbulence, channels coupled so that correcting one can d

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning

DGX agent

arXiv:2606.00671v1 Announce Type: new Abstract: We present AXIOM, a trust-first neuro-symbolic execution architecture for natural-language mathematical reasoning. In AXIOM, the language model function

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

DGX agent

arXiv:2606.02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approac

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Bayesian Inference of Nonlinear Malaria Dynamics in Ghana via an Ensemble Markov Chain Monte Carlo Sampler

DGX agent

arXiv:2606.00783v1 Announce Type: cross Abstract: Reliable quantification of malaria dynamics in sub-Saharan Africa is hindered by short, noisy, and spatially heterogeneous surveillance records. In Gh

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Before and After Temperature: A Distributional View of Creative LLM Generation

DGX agent

arXiv:2606.01451v1 Announce Type: new Abstract: Reference-free evaluation of large language model (LLM) creativity relies on perplexity, entropy, and top-1 margin. We show that a much stronger signal

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution

DGX agent

arXiv:2606.01286v1 Announce Type: cross Abstract: The rapid progress of frontier large language models has led to widespread benchmark saturation, limiting the ability of existing datasets to differen

model-releasesarxiv-cs-ai
2 Jun 2026
Model Releases

Benchmark Dataset for Catalysis on 2D MXenes

DGX agent

arXiv:2606.00794v1 Announce Type: cross Abstract: Merging first-principles calculations with machine learning (ML), we aim to accelerate the exploration of catalytic behaviour in novel materials. We f

model-releasesarxiv-cs-lg
2 Jun 2026
Model Releases

Benchmarking Large Language Models for Cryptanalysis and Side-Channel Vulnerabilities

DGX agent

arXiv:2505.24621v3 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have transformed natural language understanding and generation, leading to extensive benchmarkin

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

DGX agent

arXiv:2606.01629v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used for long-form generation, reliably evaluating long-form outputs has become a critical challenge. L

model-releasesarxiv-cs-cl
2 Jun 2026
Model Releases

Benchmarking Local LLMs for Natural-Language-to-SQL Querying in Biopharmaceutical Manufacturing: An Empirical Benchmark on Consumer-Grade Hardware

DGX agent

arXiv:2606.01338v1 Announce Type: new Abstract: Biopharmaceutical manufacturing organizations operate under regulatory frameworks such as FDA guidance, EU Good Manufacturing Practice (GMP), and the EU

model-releasesarxiv-cs-cl
2 Jun 2026
← Previous
1…163164165166167…361
Next →