AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
26 May 2026

A Blended Likelihood Approach for Achieving Fairness Using Naive Bayes

Model ReleasesDGX agent

arXiv:2605.25228v1 Announce Type: new Abstract: Concerns about algorithmic bias and fairness have increased as artificial intelligence has been incorporated into high-stakes decision-making. Tradition

A Comprehensive Dataset for Human vs. AI Generated Text Detection

Model ReleasesDGX agent

arXiv:2510.22874v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authentic

A computational phase transition for learning-to-sample from Ising models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.24752v1 Announce Type: new Abstract: We study learning-to-sample -- a basic algorithmic task underlying generative modeling -- for Ising models, a standard testbed for algorithmic ideas in

A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.25502v1 Announce Type: cross Abstract: Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because e

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

Model ReleasesDGX agent

arXiv:2605.24045v1 Announce Type: cross Abstract: Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a p

A Learning Stability Profile for Finite-Dimensional Learning Dynamics

Model ReleasesDGX agent

arXiv:2512.21208v3 Announce Type: replace Abstract: We develop a finite-dimensional sensitivity framework for studying stability in learning systems whose states include representations, parameters, a

A lift for input-convex neural network training

Model ReleasesDGX agent

arXiv:2605.24274v1 Announce Type: new Abstract: Input-convex neural networks (ICNNs) are widely used for log-concave density estimation, convex-potential normalizing flows, optimal transport, and tran

A Matched Spectral Benchmark of Quantum Inspired Feature Maps

Model ReleasesDGX agent

arXiv:2605.24324v1 Announce Type: cross Abstract: Quantum machine learning is often motivated by the idea that quantum systems can expose useful high-dimensional structure that is difficult to access

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

Model ReleasesDGX agent

arXiv:2605.23977v1 Announce Type: new Abstract: This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS,

A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training

Model ReleasesDGX agent

arXiv:2605.24006v1 Announce Type: cross Abstract: Pipeline parallelism is a key technique for distributed training of large language models because it reduces per-device parameter and activation memor

A tool to get Claude Code-style reliability from fully local models

Model ReleasesDGX agent

Ollama exposes an Anthropic-compatible Messages endpoint , allowing developers to run powerful open-source AI models locally with no API costs and pair them with Claude Code for a capable local AI cod

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

Model ReleasesDGX agent

arXiv:2605.25652v1 Announce Type: new Abstract: Free-form legal essay evaluation in NLP treats expert inter-rater stability as a single ceiling number, and treats LLM-judge agreement with that ceiling

A World Model of Radiologist Reading for Medical Image Representation Learning

Model ReleasesDGX agent

arXiv:2605.23992v1 Announce Type: cross Abstract: Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing method

Acting on the Unseen: Communication-Free Collaborative Filtering for Decentralized Multi-Robot Task Allocation

Model ReleasesDGX agent

arXiv:2605.25584v1 Announce Type: cross Abstract: Multi-robot task allocation usually assumes some combination of communication, known task models, or a coordinator. We study the opposite extreme, a r

Action-Prior Denoising for Smooth Real-Time Chunking

Model ReleasesDGX agent

arXiv:2605.25537v1 Announce Type: new Abstract: Real-time chunking (RTC) lets chunked action policies operate under inference delay by conditioning a newly generated action chunk on actions already co

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.24011v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge pl

Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity Resolution

Model ReleasesDGX agent

arXiv:2605.25814v1 Announce Type: cross Abstract: Dirty entity resolution (ER), which identifies records referring to the same real-world entity from a single, messy dataset, is a fundamental task in

Advancing Graph Few-Shot Learning via In-Context Learning

Model ReleasesDGX agent

arXiv:2605.24410v1 Announce Type: new Abstract: Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

Model ReleasesDGX agent

arXiv:2605.23974v1 Announce Type: new Abstract: Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness its

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. Howev

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

Model ReleasesDGX agent

arXiv:2505.24876v2 Announce Type: replace-cross Abstract: Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understandi

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

Model ReleasesDGX agent

arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple r

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

Model ReleasesDGX agent

arXiv:2605.25901v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language des

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

Model ReleasesDGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

Model ReleasesDGX agent

arXiv:2605.25358v1 Announce Type: cross Abstract: AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refi

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

Model ReleasesDGX agent

arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, ma

AI Content Moderation in Therapy Conversations

Model ReleasesDGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

Model ReleasesDGX agent

arXiv:2602.22769v3 Announce Type: replace Abstract: Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long-horizon memory is critical

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

Model ReleasesDGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

An annoyance with Claude right now is that changes to the interface are badly documented, resulting in frustrating dead ends. For example, l…

Model ReleasesDGX agent

An annoyance with Claude right now is that changes to the interface are badly documented, resulting in frustrating dead ends. For example, learning mode is migrating to a skill. Where is that skill? T

An Effective-Rank Audit of Alignment-Induced Activation Shifts: Confound Control, Constructive Calibration, and Limits

Model ReleasesDGX agent

arXiv:2605.24583v1 Announce Type: cross Abstract: We audit alignment-induced shifts in residual-stream activations of three open-weight instruction-tuned LLMs (Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qw

An Efficient Learning Method to Connect Observables

Model ReleasesDGX agent

arXiv:2503.01684v3 Announce Type: replace-cross Abstract: Constructing fast and accurate surrogate models is a key ingredient for making robust predictions in many topics. We introduce a new model, th

An Interactive Paradigm for Deep Research

Model ReleasesDGX agent

arXiv:2605.24266v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended q

An Interpretable CF-RL-TOPSIS Fusion Model for Skills-Aware Talent Recommendation

Model ReleasesDGX agent

arXiv:2605.24155v1 Announce Type: cross Abstract: Effective skills-aware talent recommendation must balance behavioral transition patterns, trajectory-sensitive adaptation, and inspectable occupation-

AnnotateMissense: a genome-wide annotation and benchmarking framework for missense pathogenicity prediction

Model ReleasesDGX agent

arXiv:2605.24520v1 Announce Type: cross Abstract: Missense variant interpretation remains challenging because pathogenicity depends on heterogeneous evidence from population frequency, evolutionary co

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

Model ReleasesDGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

AppOmni launches Marlin AI to automate SaaS security investigation and remediation

Model ReleasesDGX agent

Software-as-a-service security company AppOmni Inc. today launched Marlin AI, a new autonomous artificial intelligence capability built into its platform that handles correlation, investigation and gu

ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views

Model ReleasesDGX agent

arXiv:2605.24304v1 Announce Type: cross Abstract: Articulated object reconstruction from sparse-view images is an ill-posed problem that requires simultaneous inference of geometry and underlying arti

Asking LLMs to Verify First is Almost Free Lunch

Model ReleasesDGX agent

arXiv:2511.21734v2 Announce Type: replace-cross Abstract: To enhance the reasoning capabilities of Large Language Models (LLMs) without high costs of training, nor extensive test-time sampling, we int

AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

Model ReleasesDGX agent

arXiv:2605.24573v1 Announce Type: new Abstract: Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth o

Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework

Model ReleasesDGX agent

arXiv:2605.23921v1 Announce Type: cross Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questi

AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora

Model ReleasesDGX agent

arXiv:2605.25382v1 Announce Type: new Abstract: Evidence construction systems--chunk retrieval, agent memory, knowledge-graph traversal, and thematic indexing--are evaluated on separate benchmarks wit

Automated Benchmark Auditing for AI Agents and Large Language Models

Model ReleasesDGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

Autoregression-Free Neural Operators for Time-Dependent PDEs

Model ReleasesDGX agent

arXiv:2605.25413v1 Announce Type: cross Abstract: Neural operators learn mappings from function-dependent inputs to solutions, providing an effective framework for solving partial differential equatio

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

Model ReleasesDGX agent

arXiv:2605.24183v1 Announce Type: cross Abstract: We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through latent world recovery. AvalancheBench improves on existing benc

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

Model ReleasesDGX agent

arXiv:2605.24652v1 Announce Type: new Abstract: Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios inv

Axis-Aligned Semantics for ODRL: Resolving Dimensional Ambiguity in Policy Constraints

Model ReleasesDGX agent

arXiv:2602.19878v3 Announce Type: replace Abstract: The Open Digital Rights Language (ODRL) represents policy constraints as triples of a left operand, an operator, and a value. Several spatial operan

Bayesian Distributional Models of Executive Functioning

Model ReleasesDGX agent

arXiv:2510.00387v3 Announce Type: replace Abstract: This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian

BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data

Model ReleasesDGX agent

arXiv:2605.25549v1 Announce Type: cross Abstract: High-quality expert chain-of-thought (CoT) data is one of the core bottlenecks in large language model (LLM) post-training. Existing data production m

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

Model ReleasesDGX agent

arXiv:2605.21602v2 Announce Type: replace Abstract: Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response pat

Benchmarking and Learning Real-World Customer Service Dialogue

Model ReleasesDGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

Model ReleasesDGX agent

arXiv:2605.24297v1 Announce Type: cross Abstract: Which fine-tuning signals improve patent embedding models, and do gains transfer across patent landscapes? We benchmark 22 embedding models, from 22M-

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

Model ReleasesDGX agent

arXiv:2605.25764v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

Model ReleasesDGX agent

arXiv:2605.24423v1 Announce Type: new Abstract: In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

Model ReleasesDGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction

Model ReleasesDGX agent

arXiv:2605.24657v1 Announce Type: new Abstract: Major LLM platforms deploy models in an inference-only configuration: the model serves requests but never updates per-user weights. Users must repeatedl

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

Model ReleasesDGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching

Model ReleasesDGX agent

arXiv:2605.25558v1 Announce Type: new Abstract: Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Curre

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

Model ReleasesDGX agent

arXiv:2605.26100v1 Announce Type: cross Abstract: Code review is a critical practice in software engineering, yet the growing scale and frequency of code patches in modern projects, together with the

BODHI: Precise OS Kernel Specification Inference

Model ReleasesDGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

← Previous
1…209210211212213…377
Next →