AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
All
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,620 results
Model Releases

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

DGX agent

arXiv:2605.24011v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge pl

model-releasesarxiv-cs-ai
26 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity Resolution

DGX agent

arXiv:2605.25814v1 Announce Type: cross Abstract: Dirty entity resolution (ER), which identifies records referring to the same real-world entity from a single, messy dataset, is a fundamental task in

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Advancing Graph Few-Shot Learning via In-Context Learning

DGX agent

arXiv:2605.24410v1 Announce Type: new Abstract: Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

DGX agent

arXiv:2605.23974v1 Announce Type: new Abstract: Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness its

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

DGX agent

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. Howev

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

DGX agent

arXiv:2505.24876v2 Announce Type: replace-cross Abstract: Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understandi

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

DGX agent

arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple r

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

DGX agent

arXiv:2605.25901v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language des

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

DGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

DGX agent

arXiv:2605.25358v1 Announce Type: cross Abstract: AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

DGX agent

arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, ma

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI Content Moderation in Therapy Conversations

DGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

DGX agent

arXiv:2602.22769v3 Announce Type: replace Abstract: Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long-horizon memory is critical

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

DGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

An annoyance with Claude right now is that changes to the interface are badly documented, resulting in frustrating dead ends. For example, l…

DGX agent

An annoyance with Claude right now is that changes to the interface are badly documented, resulting in frustrating dead ends. For example, learning mode is migrating to a skill. Where is that skill? T

model-releasesethan-mollick--x
26 May 2026
Model Releases

An Effective-Rank Audit of Alignment-Induced Activation Shifts: Confound Control, Constructive Calibration, and Limits

DGX agent

arXiv:2605.24583v1 Announce Type: cross Abstract: We audit alignment-induced shifts in residual-stream activations of three open-weight instruction-tuned LLMs (Llama-3.1-8B-Instruct, Gemma-2-9B-it, Qw

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

An Efficient Learning Method to Connect Observables

DGX agent

arXiv:2503.01684v3 Announce Type: replace-cross Abstract: Constructing fast and accurate surrogate models is a key ingredient for making robust predictions in many topics. We introduce a new model, th

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

An Interactive Paradigm for Deep Research

DGX agent

arXiv:2605.24266v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended q

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

An Interpretable CF-RL-TOPSIS Fusion Model for Skills-Aware Talent Recommendation

DGX agent

arXiv:2605.24155v1 Announce Type: cross Abstract: Effective skills-aware talent recommendation must balance behavioral transition patterns, trajectory-sensitive adaptation, and inspectable occupation-

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AnnotateMissense: a genome-wide annotation and benchmarking framework for missense pathogenicity prediction

DGX agent

arXiv:2605.24520v1 Announce Type: cross Abstract: Missense variant interpretation remains challenging because pathogenicity depends on heterogeneous evidence from population frequency, evolutionary co

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Anticipate and Learn: Unleashing Idle-Time Compute in Proactive Agents

DGX agent

arXiv:2605.25971v1 Announce Type: new Abstract: While AI agents demonstrate remarkable capabilities in reasoning and tool use, they remain fundamentally reactive: they compute responses only after exp

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AppOmni launches Marlin AI to automate SaaS security investigation and remediation

DGX agent

Software-as-a-service security company AppOmni Inc. today launched Marlin AI, a new autonomous artificial intelligence capability built into its platform that handles correlation, investigation and gu

model-releasessiliconangle
26 May 2026
Model Releases

ArtSplat: Feed-Forward Articulated 3D Gaussian Splatting from Sparse Multi-State Uncalibrated Views

DGX agent

arXiv:2605.24304v1 Announce Type: cross Abstract: Articulated object reconstruction from sparse-view images is an ill-posed problem that requires simultaneous inference of geometry and underlying arti

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Asking LLMs to Verify First is Almost Free Lunch

DGX agent

arXiv:2511.21734v2 Announce Type: replace-cross Abstract: To enhance the reasoning capabilities of Large Language Models (LLMs) without high costs of training, nor extensive test-time sampling, we int

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AstroMind: A High-Fidelity Benchmark for Spacecraft Behavior Reasoning Based on Large Language Models

DGX agent

arXiv:2605.24573v1 Announce Type: new Abstract: Understanding why a spacecraft maneuvers -- rather than simply that it did -- is an increasingly important problem for space domain awareness as Earth o

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Authority Signals in Claude AI Health Citations: A Descriptive Analysis Using the Authority Signals Framework

DGX agent

arXiv:2605.23921v1 Announce Type: cross Abstract: This study seeks to determine the authority signals used by Anthropic's Claude AI in its presentation of sources when answering consumer health questi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora

DGX agent

arXiv:2605.25382v1 Announce Type: new Abstract: Evidence construction systems--chunk retrieval, agent memory, knowledge-graph traversal, and thematic indexing--are evaluated on separate benchmarks wit

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Automated Benchmark Auditing for AI Agents and Large Language Models

DGX agent

arXiv:2605.26079v1 Announce Type: new Abstract: Modern AI benchmarks operate at a complexity that outpaces traditional verification methods. Tasks authored by domain experts often contain implicit ass

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Autoregression-Free Neural Operators for Time-Dependent PDEs

DGX agent

arXiv:2605.25413v1 Announce Type: cross Abstract: Neural operators learn mappings from function-dependent inputs to solutions, providing an effective framework for solving partial differential equatio

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AvalancheBench: Evaluating Enterprise Data Agents Through Latent World Recovery

DGX agent

arXiv:2605.24183v1 Announce Type: cross Abstract: We introduce AvalancheBench, a benchmark for evaluating enterprise data agents through latent world recovery. AvalancheBench improves on existing benc

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models

DGX agent

arXiv:2605.24652v1 Announce Type: new Abstract: Rapid advances in audio-video (AV) generation have enabled high-fidelity synthesis with synchronized sound, particularly for human-related scenarios inv

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Axis-Aligned Semantics for ODRL: Resolving Dimensional Ambiguity in Policy Constraints

DGX agent

arXiv:2602.19878v3 Announce Type: replace Abstract: The Open Digital Rights Language (ODRL) represents policy constraints as triples of a left operand, an operator, and a value. Several spatial operan

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Bayesian Distributional Models of Executive Functioning

DGX agent

arXiv:2510.00387v3 Announce Type: replace Abstract: This study uses controlled simulations with known ground-truth parameters to evaluate how Distributional Latent Variable Models (DLVM) and Bayesian

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

BC Protocol: Structured Dual-Expert Dialogue for Eliciting High-Quality Chain-of-Thought Post-Training Data

DGX agent

arXiv:2605.25549v1 Announce Type: cross Abstract: High-quality expert chain-of-thought (CoT) data is one of the core bottlenecks in large language model (LLM) post-training. Existing data production m

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs

DGX agent

arXiv:2605.21602v2 Announce Type: replace Abstract: Many safety and alignment failures of large language models (LLMs) occur due to out-of-distribution (OOD) situations: unusual prompt or response pat

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking and Learning Real-World Customer Service Dialogue

DGX agent

arXiv:2510.22143v3 Announce Type: replace Abstract: Existing benchmarks and training pipelines for industrial intelligent customer service (ICS) remain misaligned with real-world dialogue requirements

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Benchmarking Patent Embeddings: A Multi-Task Evaluation of 22 Models Across Retrieval, Classification, and Clustering

DGX agent

arXiv:2605.24297v1 Announce Type: cross Abstract: Which fine-tuning signals improve patent embedding models, and do gains transfer across patent landscapes? We benchmark 22 embedding models, from 22M-

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking Pathology Foundation Models for Spatial Domain Understanding

DGX agent

arXiv:2605.25764v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have emerged as a core approach for learning transferable representations from whole slide images (WSIs), and they

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Benchmarking the Limits of In-Context Reinforcement Learning for Ad-Hoc Teamwork

DGX agent

arXiv:2605.24423v1 Announce Type: new Abstract: In-Context Reinforcement Learning (ICRL) has enabled foundation agents to adapt instantaneously to novel tasks, yet its efficacy in Ad-Hoc Teamwork (AHT

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

DGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Inference-Only Deployment: Comparing Weight-Based Consolidation Against Cascading Compaction

DGX agent

arXiv:2605.24657v1 Announce Type: new Abstract: Major LLM platforms deploy models in an inference-only configuration: the model serves requests but never updates per-user weights. Users must repeatedl

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

DGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching

DGX agent

arXiv:2605.25558v1 Announce Type: new Abstract: Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Curre

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

DGX agent

arXiv:2605.26100v1 Announce Type: cross Abstract: Code review is a critical practice in software engineering, yet the growing scale and frequency of code patches in modern projects, together with the

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

BODHI: Precise OS Kernel Specification Inference

DGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training

DGX agent

arXiv:2509.24050v4 Announce Type: replace Abstract: Device-cloud collaboration holds promise for deploying large language models (LLMs), leveraging lightweight on-device models for efficiency while re

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Building an Adversarial Malware Dataset by Family and Type: Generation, Evasion, and Poisoning Evaluation

DGX agent

arXiv:2605.25937v1 Announce Type: cross Abstract: We present a dataset of adversarial malware samples derived from the public RawMal-TF collection of real-world malware binaries. Using a suite of adve

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

DGX agent

arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint

model-releasesarxiv-cs-ai
26 May 2026
← Previous
1…261262263264265…472
Next →