AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,603 results
Model Releases

Cattle Trade: A Multi-Agent Benchmark for LLM Bluffing, Bidding, and Bargaining

DGX agent

arXiv:2605.14537v1 Announce Type: new Abstract: We introduce extsc{Cattle Trade, a multi-agent benchmark for evaluating large language models (LLMs) as agents in strategic reasoning under imperfect in

model-releasesarxiv-cs-ai
15 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

CausalReasoningBenchmark: A Real-World Benchmark for Disentangled Evaluation of Causal Identification and Estimation

DGX agent

arXiv:2602.20571v2 Announce Type: replace Abstract: Many benchmarks for automated causal inference evaluate a system's performance based on a single numerical output, such as an Average Treatment Effe

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA

DGX agent

arXiv:2605.14928v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) have achieved impressive results on standard image-text tasks, yet their potential for visual procedure

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization

DGX agent

arXiv:2511.15408v2 Announce Type: replace-cross Abstract: Chinese demonstrates high semantic compactness and rich metaphorical expressiveness, enabling limited text to convey dense meanings while incr

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

CineMesh4D: Personalized 4D Whole Heart Reconstruction from Sparse Cine MRI

DGX agent

arXiv:2605.13994v1 Announce Type: cross Abstract: Accurate 3D+t whole-heart mesh reconstruction from cine MRI is a clinically crucial yet technically challenging task. The difficulty of this task aris

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Claude Code's product lead talks usage limits, transparency, and the 'lean harness'

DGX agent

Claude Code's product lead addresses how Anthropic tunes the 'harness' (the structural layer around the model) for each new model release to optimize performance and reduce verbosity. The company comm

model-releasesars-technica
15 May 2026
Model Releases

ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents

DGX agent

arXiv:2605.14133v1 Announce Type: new Abstract: Interactive agent benchmarks face a tension between scalable construction and realistic workflow evaluation. Hand-authored tasks are expensive to extend

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

CLOVER: Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning

DGX agent

arXiv:2605.15120v1 Announce Type: cross Abstract: End-to-end autonomous driving planners are commonly trained by imitating a single logged trajectory, yet evaluated by rule-based planning metrics that

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

CoCoEdit: Content-Consistent Image Editing via Region Regularized Reinforcement Learning

DGX agent

arXiv:2602.14068v2 Announce Type: replace Abstract: Image editing has achieved impressive results with the development of large-scale generative models. However, existing models mainly focus on the ed

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Cognitive-Uncertainty Guided Knowledge Distillation for Accurate Classification of Student Misconceptions

DGX agent

arXiv:2605.14752v1 Announce Type: cross Abstract: Accurately identifying student misconceptions is crucial for personalized education but faces three challenges: (1) data scarcity with long-tail distr

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

DGX agent

arXiv:2605.13950v1 Announce Type: cross Abstract: Autonomous language-model agents are increasingly evaluated on long-horizon tool-use tasks, but existing benchmarks rarely capture the complexity and

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Communication-Efficient Federated Fine-Tuning

DGX agent

arXiv:2505.04535v3 Announce Type: replace Abstract: Federated Learning (FL) enables the utilization of vast, previously inaccessible data sources. At the same time, pre-trained Language Models (LMs) h

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

Correctness-Aware Repository Filtering Under Maximum Effective Context Window Constraints

DGX agent

arXiv:2605.14362v1 Announce Type: cross Abstract: Context window efficiency is a practical constraint in large language model (LLM)-based developer tools. Paulsen [12] shows that all tested models deg

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering

DGX agent

arXiv:2506.08584v4 Announce Type: replace Abstract: Medical question answering (QA) benchmarks often focus on multiple-choice or fact-based tasks, leaving open-ended answers to real patient questions

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

CRANE: Constrained Reasoning Injection for Code Agents via Nullspace Editing

DGX agent

arXiv:2605.14084v1 Announce Type: cross Abstract: Code agents must both reason over long-horizon repository state and obey strict tool-use protocols. In paired Instruct/Thinking checkpoints, these cap

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Critic-Driven Voronoi-Quantization for Distilling Deep RL Policies to Explainable Models

DGX agent

arXiv:2605.14897v1 Announce Type: cross Abstract: Despite many successful attempts at explaining Deep Reinforcement Learning policies using distillation, it remains difficult to balance the performanc

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

CUICurate: A GraphRAG-based Framework for Automated Clinical Concept Curation for NLP applications

DGX agent

arXiv:2602.17949v2 Announce Type: replace-cross Abstract: Background: Clinical named entity recognition tools commonly map free text to Unified Medical Language System (UMLS) Concept Unique Identifier

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

CurveBench: A Benchmark for Exact Topological Reasoning over Nested Jordan Curves

DGX agent

arXiv:2605.14068v1 Announce Type: new Abstract: We introduce CurveBench, a benchmark for hierarchical topological reasoning from visual input. CurveBench consists of extbf{756 images} of pairwise non-

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning

DGX agent

arXiv:2605.14386v1 Announce Type: cross Abstract: We present Darwin Family, a framework for training-free evolutionary merging of large language models via gradient-free weight-space recombination. We

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

DGX agent

arXiv:2605.14379v1 Announce Type: cross Abstract: Finding approximate equilibria for large-scale imperfect-information competitive games such as StarCraft, Dota, and CounterStrike remains computationa

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Databricks brings GPT-5.5 to enterprise agent workflows

DGX agent

Databricks has integrated OpenAI's GPT-5.5 model into enterprise agent workflows, enabling organizations to build and deploy AI agents with advanced language capabilities. This partnership leverages D

model-releasesopenai
15 May 2026
Model Releases

Dataiku launches governed AI workflow builder inside Snowflake

DGX agent

Dataiku Inc. today announced a deeper integration with Snowflake Inc. aimed at simplifying the creation of enterprise AI agents while preserving governance and operational controls, a growing concern

model-releasessiliconangle
15 May 2026
Model Releases

Deceive, Detect, and Disclose: Large Language Models Play Mini-Mafia

DGX agent

arXiv:2509.23023v3 Announce Type: replace Abstract: Large language models are increasingly deployed in multi-agent settings whose outcomes hinge on social intelligence, motivating evaluations of their

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Deep Image Segmentation via Discriminant Feature Learning

DGX agent

arXiv:2605.14609v1 Announce Type: new Abstract: Accurate image segmentation remains challenging, particularly in generating sharp, confident boundaries. While modern architectures have advanced the fi

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Denoising-GS: Gaussian Splatting with Spatial-aware Denoising

DGX agent

arXiv:2605.14880v1 Announce Type: new Abstract: Recent advances in 3D Gaussian Splatting (3DGS) have achieved remarkable success in high-fidelity Novel View Synthesis (NVS), yet the optimization proce

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Derivation Prompting: A Logic-Based Method for Improving Retrieval-Augmented Generation

DGX agent

arXiv:2605.14053v1 Announce Type: cross Abstract: The application of Large Language Models to Question Answering has shown great promise, but important challenges such as hallucinations and erroneous

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA)

DGX agent

arXiv:2511.13397v2 Announce Type: replace-cross Abstract: The remarkable progress of Vision-Language Models (VLMs) on a variety of tasks has raised interest in their application to automated driving.

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Discovering Physical Directions in Weight Space: Composing Neural PDE Experts

DGX agent

arXiv:2605.14546v1 Announce Type: new Abstract: Recent advances in neural operators have made partial differential equation (PDE) surrogate modeling increasingly scalable and transferable through larg

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

Distribution-Aware Algorithm Design with LLM Agents

DGX agent

arXiv:2605.14141v1 Announce Type: new Abstract: We study learning when the learned object is executable solver code rather than a predictor. In this setting, correctness is not enough: two solvers may

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Do Coding Agents Understand Least-Privilege Authorization?

DGX agent

arXiv:2605.14859v1 Announce Type: cross Abstract: As coding agents gain access to shells, repositories, and user files, least-privilege authorization becomes a prerequisite for safe deployment: an age

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Do-Undo Bench: Reversibility for Action Understanding in Image Generation

DGX agent

arXiv:2512.13609v2 Announce Type: replace Abstract: We introduce the Do-Undo task and benchmark to address a critical gap in vision-language models: understanding and generating plausible scene transf

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

'do you want meet in the Devin train car?' 'oh actually I'm in the Langsmith car near the front' 'Let's just meet in front of the Mistral wa…

DGX agent

This appears to be a humorous exchange from X/Twitter creator Harrison Chase referencing train cars named after AI/ML tools and services (Devin, Langsmith, and Mistral), likely making a joke about the

model-releasesharrison-chase--x
15 May 2026
Model Releases

Does RAG Know When Retrieval Is Wrong? Diagnosing Context Compliance under Knowledge Conflict

DGX agent

arXiv:2605.14473v1 Announce Type: cross Abstract: The Context-Compliance Regime in Retrieval-Augmented Generation (RAG) occurs when retrieved context dominates the final answer even when it conflicts

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

DVMap: Fine-Grained Pluralistic Value Alignment via High-Consensus Demographic-Value Mapping

DGX agent

arXiv:2605.14420v1 Announce Type: new Abstract: Current Large Language Models (LLMs) typically rely on coarse-grained national labels for pluralistic value alignment. However, such macro-level supervi

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis

DGX agent

arXiv:2605.14842v1 Announce Type: new Abstract: Humans naturally communicate through abstract concepts like 'mood'. However, current image editing benchmarks focus primarily on explicit, literal comma

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Efficient Online Conformal Selection with Limited Feedback

DGX agent

arXiv:2605.14953v1 Announce Type: new Abstract: We address the problem of conformal selection, where an agent must select a minimal subset of options to ensure that at least one ``success'' is identif

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

Elastic Spiking Transformers for Efficient Gesture Understanding

DGX agent

arXiv:2605.13869v1 Announce Type: cross Abstract: Spiking Neural Networks (SNNs), particularly Spiking Transformers, offer energy-efficient processing of event-based sensor data for healthcare applica

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring

DGX agent

arXiv:2605.14589v1 Announce Type: new Abstract: Extending the context window of large language models typically requires training on sequences at the target length, incurring quadratic memory and comp

model-releasesarxiv-cs-cl
15 May 2026
Model Releases

Enhancing Few-Shot Classification of Benchmark and Disaster Imagery with ABHFA-Net

DGX agent

arXiv:2510.18326v3 Announce Type: replace Abstract: The rising incidence of natural and human-induced disasters necessitates robust visual recognition systems capable of operating under limited labele

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation

DGX agent

arXiv:2605.15199v1 Announce Type: cross Abstract: Multi-shot video generation extends single-shot generation to coherent visual narratives, yet maintaining consistent characters, objects, and location

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Eradicating Negative Transfer in Multi-Physics Foundation Models via Sparse Mixture-of-Experts Routing

DGX agent

arXiv:2605.15179v1 Announce Type: cross Abstract: Scaling Scientific Machine Learning (SciML) toward universal foundation models is bottlenecked by negative transfer: the simultaneous co-training of d

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

EvolveMem:Self-Evolving Memory Architecture via AutoResearch for LLM Agents

DGX agent

arXiv:2605.13941v1 Announce Type: cross Abstract: Long-term memory is essential for LLM agents that operate across multiple sessions, yet existing memory systems treat retrieval infrastructure as fixe

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Exemplar Partitioning for Mechanistic Interpretability

DGX agent

arXiv:2605.14347v1 Announce Type: new Abstract: We introduce Exemplar Partitioning (EP), an unsupervised method for constructing interpretable feature dictionaries from large language model activation

model-releasesarxiv-cs-lg
15 May 2026
Model Releases

ExploitBench: A Capability Ladder Benchmark for LLM Cybersecurity Agents

DGX agent

arXiv:2605.14153v1 Announce Type: cross Abstract: Exploitation is not a binary event. It is a ladder of acquiring progressive capabilities, from executing a single buggy line of code to taking full co

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

Exploring Vision-Language Models for Online Signature Verification: A Zero-Shot Capability Study

DGX agent

arXiv:2605.14845v1 Announce Type: new Abstract: Recent advancements in Vision-Language Models (VLMs) have demonstrated strong capabilities in general visual reasoning, yet their applicability to rigor

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

Fast Rates for Inverse Reinforcement Learning

DGX agent

arXiv:2605.14599v1 Announce Type: cross Abstract: We establish novel structural and statistical results for entropy-regularized min-max inverse reinforcement learning (Min-Max-IRL) with linear reward

model-releasesarxiv-cs-ai
15 May 2026
Model Releases

FedStain: Modeling Higher-Order Stain Statistics for Federated Domain Generalization in Computational Pathology

DGX agent

arXiv:2605.14590v1 Announce Type: new Abstract: Robust whole-slide image (WSI) analysis under strict data-governance remains challenging due to substantial cross-institutional stain heterogeneity. Dom

model-releasesarxiv-cs-cv
15 May 2026
Model Releases

FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching

DGX agent

arXiv:2604.06757v2 Announce Type: replace Abstract: Multimodal generation has long been dominated by text-driven pipelines where language dictates vision but cannot reason or create within it. We chal

model-releasesarxiv-cs-cv
15 May 2026
← Previous
1…309310311312313…471
Next →