AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

DGX agent

arXiv:2605.25901v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language des

model-releasesarxiv-cs-ro
26 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

DGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

DGX agent

arXiv:2605.25358v1 Announce Type: cross Abstract: AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

DGX agent

arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, ma

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI Content Moderation in Therapy Conversations

DGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

DGX agent

arXiv:2602.22769v3 Announce Type: replace Abstract: Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long-horizon memory is critical

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

DGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

An Effective-Rank Audit of Alignment-Induced Activation Shifts: Confound Control, Constructive Calibration, and Limits

DGX agent

arXiv:2605.24583v1 Announce Type: cross Abstract: We audit alignment-induced shifts in residual-stream activations of three open-weight instruction-tuned LLMs (Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qw

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

An Efficient Learning Method to Connect Observables

DGX agent

arXiv:2503.01684v3 Announce Type: replace-cross Abstract: Constructing fast and accurate surrogate models is a key ingredient for making robust predictions in many topics. We introduce a new model, th

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

An Interactive Paradigm for Deep Research

DGX agent

arXiv:2605.24266v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended q

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

An Interpretable CF-RL-TOPSIS Fusion Model for Skills-Aware Talent Recommendation

DGX agent

arXiv:2605.24155v1 Announce Type: cross Abstract: Effective skills-aware talent recommendation must balance behavioral transition patterns, trajectory-sensitive adaptation, and inspectable occupation-

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AnnotateMissense: a genome-wide annotation and benchmarking framework for missense pathogenicity prediction

DGX agent

arXiv:2605.24520v1 Announce Type: cross Abstract: Missense variant interpretation remains challenging because pathogenicity depends on heterogeneous evidence from population frequency, evolutionary co

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

DGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views

DGX agent

arXiv:2605.24304v1 Announce Type: cross Abstract: Articulated object reconstruction from sparse-view images is an ill-posed problem that requires simultaneous inference of geometry and underlying arti

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Asking LLMs to Verify First is Almost Free Lunch

DGX agent

arXiv:2511.21734v2 Announce Type: replace-cross Abstract: To enhance the reasoning capabilities of Large Language Models (LLMs) without high costs of training, nor extensive test-time sampling, we int

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

DGX agent

arXiv:2605.24573v1 Announce Type: new Abstract: Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth o

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework

DGX agent

arXiv:2605.23921v1 Announce Type: cross Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora

DGX agent

arXiv:2605.25382v1 Announce Type: new Abstract: Evidence construction systems--chunk retrieval, agent memory, knowledge-graph traversal, and thematic indexing--are evaluated on separate benchmarks wit

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Automated Benchmark Auditing for AI Agents and Large Language Models

DGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Autoregression-Free Neural Operators for Time-Dependent PDEs

DGX agent

arXiv:2605.25413v1 Announce Type: cross Abstract: Neural operators learn mappings from function-dependent inputs to solutions, providing an effective framework for solving partial differential equatio

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

DGX agent

arXiv:2605.24183v1 Announce Type: cross Abstract: We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through latent world recovery. AvalancheBench improves on existing benc

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

DGX agent

arXiv:2605.24652v1 Announce Type: new Abstract: Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios inv

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Axis-Aligned Semantics for ODRL: Resolving Dimensional Ambiguity in Policy Constraints

DGX agent

arXiv:2602.19878v3 Announce Type: replace Abstract: The Open Digital Rights Language (ODRL) represents policy constraints as triples of a left operand, an operator, and a value. Several spatial operan

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Bayesian Distributional Models of Executive Functioning

DGX agent

arXiv:2510.00387v3 Announce Type: replace Abstract: This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data

DGX agent

arXiv:2605.25549v1 Announce Type: cross Abstract: High-quality expert chain-of-thought (CoT) data is one of the core bottlenecks in large language model (LLM) post-training. Existing data production m

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

DGX agent

arXiv:2605.21602v2 Announce Type: replace Abstract: Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response pat

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking and Learning Real-World Customer Service Dialogue

DGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

DGX agent

arXiv:2605.24297v1 Announce Type: cross Abstract: Which fine-tuning signals improve patent embedding models, and do gains transfer across patent landscapes? We benchmark 22 embedding models, from 22M-

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

DGX agent

arXiv:2605.25764v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

DGX agent

arXiv:2605.24423v1 Announce Type: new Abstract: In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

DGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction

DGX agent

arXiv:2605.24657v1 Announce Type: new Abstract: Major LLM platforms deploy models in an inference-only configuration: the model serves requests but never updates per-user weights. Users must repeatedl

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

DGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching

DGX agent

arXiv:2605.25558v1 Announce Type: new Abstract: Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Curre

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

DGX agent

arXiv:2605.26100v1 Announce Type: cross Abstract: Code review is a critical practice in software engineering, yet the growing scale and frequency of code patches in modern projects, together with the

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

BODHI: Precise OS Kernel Specification Inference

DGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training

DGX agent

arXiv:2509.24050v4 Announce Type: replace Abstract: Device-cloud collaboration holds promise for deploying large language models (LLMs), leveraging lightweight on-device models for efficiency while re

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Building an Adversarial Malware Dataset by Family and Type: Generation, Evasion, and Poisoning Evaluation

DGX agent

arXiv:2605.25937v1 Announce Type: cross Abstract: We present a dataset of adversarial malware samples derived from the public RawMal-TF collection of real-world malware binaries. Using a suite of adve

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

DGX agent

arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?

DGX agent

arXiv:2605.23913v1 Announce Type: cross Abstract: Cloud-hosted large language models (LLMs) commonly rely on LoRA for domain adaptation, yet domain data are distributed across multiple edge devices an

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Cascade-KDE: Robust Time-Series Restoration under Out-of-Distribution Impulse Corruptions

DGX agent

arXiv:2605.24055v1 Announce Type: cross Abstract: Real-world time-series data in industrial sensing, healthcare, and energy systems is often corrupted by a mixture of Gaussian noise and occasional lar

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

DGX agent

arXiv:2605.25891v1 Announce Type: cross Abstract: We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

DGX agent

arXiv:2605.26029v1 Announce Type: new Abstract: We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates bo

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Chain-of-Thought Hijacking

DGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ChainLearn: A Blockchain-Based Capacity-Aware Framework for Federated Ensemble Learning

DGX agent

arXiv:2605.24418v1 Announce Type: new Abstract: Federated learning is used in medical imaging where privacy prohibits centralizing data. Standard federated algorithms assume homogeneous hardware, iden

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale

DGX agent

arXiv:2605.24305v1 Announce Type: cross Abstract: Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

DGX agent

arXiv:2510.02361v2 Announce Type: replace-cross Abstract: Transformer-based large models excel in natural language processing and computer vision, but face severe computational inefficiencies due to t

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

DGX agent

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundat

model-releasesarxiv-cs-ai
26 May 2026
← Previous
1…197198199200201…361
Next →