AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,797 results
Model Releases

VitaBench 2.0: Evaluating Personalized and Proactive Agents in Long-Term User Interactions

DGX agent

arXiv:2605.27141v1 Announce Type: new Abstract: Large language models (LLMs) have evolved into interactive agents that collaborate with users in real-world tasks. Effective collaboration in such setti

model-releasesarxiv-cs-ai
27 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Warp’s big bet on building open source with GPT-5.5

DGX agent

Warp is making a significant investment in developing open source tools and integrations built on GPT-5.5, OpenAI's advanced language model. The initiative aims to leverage GPT-5.5's capabilities to c

model-releasesopenai
27 May 2026
Model Releases

We’ve been putting a lot of effort into making Claude Code more responsive & reliable. Here’s an update on everything we’ve done:

DGX agent

Anthropic has made improvements to Claude Code's responsiveness and reliability, with Thariq providing details on the enhancements implemented. The update likely covers performance optimizations, bug

model-releasesthariq--x
27 May 2026
Model Releases

We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf,…

DGX agent

We've created the world's fastest PDF parser ⚡️ And it's more accurate than any other open-source, model-free PDF parser out there (pymupdf, pypdf, markitdown, pdftotext, opendataloader, pymupdf4llm)

model-releasesjerry-liu--x
27 May 2026
Model Releases

What Demands Attention in Urban Street Scenes? From Scene Understanding towards Road Safety: A Survey of Vision-driven Datasets and Studies

DGX agent

arXiv:2507.06513v3 Announce Type: replace Abstract: Advances in vision-based sensors and computer vision algorithms have significantly improved the analysis and understanding of traffic scenarios. To

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

DGX agent

arXiv:2605.26795v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly und

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

What Molecular Structure Cannot Tell Us: A Taxonomy of Explainability Gaps in GNN-Based Drug Toxicity Prediction

DGX agent

arXiv:2605.26183v1 Announce Type: cross Abstract: Graph Neural Networks (GNNs) have emerged as a structurally natural approach for molecular toxicity prediction, operating directly on atomic connectiv

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

When Does Deep RL Beat Calibrated Baselines? A Benchmark Study on Adaptive Resource Control

DGX agent

arXiv:2605.26418v1 Announce Type: cross Abstract: A properly calibrated rule-based autoscaler can beat every one of six mainstream deep reinforcement learning (DRL) algorithms on cost across every wor

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation

DGX agent

arXiv:2509.26600v2 Announce Type: replace-cross Abstract: As LLMs rapidly saturate existing benchmarks, automated benchmark creation using LLMs (LLM-as-a-benchmark) -- where a model generates test inp

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Why Prompt Optimization Works, and Why It Sometimes Doesn't: A Causal-Inspired Edit-Level Analysis

DGX agent

arXiv:2605.26655v1 Announce Type: new Abstract: Automated prompt optimization methods (e.g., DSpy, TextGrad) can substantially improve the performance of large language model (LLM), however, their gen

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

DGX agent

arXiv:2605.26660v1 Announce Type: new Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Workflow Closure Is Not Scientific Closure in Auto-Research Systems

DGX agent

arXiv:2605.26200v1 Announce Type: cross Abstract: This paper argues that workflow closure is not scientific closure in auto-research systems. Current systems can increasingly complete research-like lo

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Xreal launches a new sub-brand, X by Xreal, and the $299 a01 display glasses with micro OLED displays, 50° FOV, and 62g weight, set for July release in the US (Scott Stein/CNET)

DGX agent

Scott Stein / CNET: Xreal launches a new sub-brand, X by Xreal, and the 299 a01 display glasses with micro OLED displays, 50° FOV, and 62g weight, set for July release in the US — X by Xreal is arrivi

model-releasestechmeme
27 May 2026
Model Releases

Your Agents Are Aging Too: Agent Lifespan Engineering for Deployed Systems

DGX agent

arXiv:2605.26302v1 Announce Type: new Abstract: Long-lived AI agents are increasingly deployed as persistent operational systems, yet they are still evaluated like freshly initialized models. Day-one

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

// Your Agents are Aging Too // Huh!? They need 'sleep,' and now they are aging? Joke aside, great write-up on reliable agentic engineering.…

DGX agent

// Your Agents are Aging Too // Huh!? They need 'sleep,' and now they are aging? Joke aside, great write-up on reliable agentic engineering. This new research introduces AgingBench, a longitudinal rel

model-releasesdair-ai--x
27 May 2026
Model Releases

Zero-Shot MARL Benchmark in the Cyber-Physical Mobility Lab

DGX agent

arXiv:2601.16578v2 Announce Type: replace Abstract: We present a reproducible benchmark for evaluating sim-to-real transfer of Multi-Agent Reinforcement Learning (MARL) policies for Connected and Auto

model-releasesarxiv-cs-ro
27 May 2026
Model Releases

Zero-Shot Object Re-Identification in Egocentric Kitchen Videos via Multi-Stage SAM3 Feature Fusion

DGX agent

arXiv:2605.26383v1 Announce Type: new Abstract: Object re-identification (ReID) in egocentric kitchen videos is challenging due to rapid viewpoint changes, frequent occlusions, cluttered scenes, and l

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

7AI launches PLAID ELITE fully managed agentic security operations service

DGX agent

Agentic artificial intelligence security startup 7AI Inc. today announced the launch of PLAID ELITE, a fully managed AI-native security operations service. The new service combines autonomous investig

model-releasessiliconangle
26 May 2026
Model Releases

A Blended Likelihood Approach for Achieving Fairness Using Naive Bayes

DGX agent

arXiv:2605.25228v1 Announce Type: new Abstract: Concerns about algorithmic bias and fairness have increased as artificial intelligence has been incorporated into high-stakes decision-making. Tradition

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A Comprehensive Dataset for Human vs. AI Generated Text Detection

DGX agent

arXiv:2510.22874v3 Announce Type: replace Abstract: The rapid advancement of large language models (LLMs) has led to increasingly human-like AI-generated text, raising concerns about content authentic

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

A computational phase transition for learning-to-sample from Ising models

DGX agent

arXiv:2605.24752v1 Announce Type: new Abstract: We study learning-to-sample -- a basic algorithmic task underlying generative modeling -- for Ising models, a standard testbed for algorithmic ideas in

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A Controlled Synthetic Benchmark for Educational Aspect-Based Sentiment Analysis

DGX agent

arXiv:2605.25502v1 Announce Type: cross Abstract: Educational aspect-based sentiment analysis (ABSA) can support course improvement, but public aspect-labeled student feedback remains scarce because e

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

A Large-Scale Dataset and Benchmark: Do Protein-Ligand Models Learn Binding Sites or Just Binding Likelihood?

DGX agent

arXiv:2605.24045v1 Announce Type: cross Abstract: Protein-ligand modeling underpins computational drug discovery and molecular design. Existing protein-ligand benchmarks typically evaluate whether a p

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

A Learning Stability Profile for Finite-Dimensional Learning Dynamics

DGX agent

arXiv:2512.21208v3 Announce Type: replace Abstract: We develop a finite-dimensional sensitivity framework for studying stability in learning systems whose states include representations, parameters, a

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A lift for input-convex neural network training

DGX agent

arXiv:2605.24274v1 Announce Type: new Abstract: Input-convex neural networks (ICNNs) are widely used for log-concave density estimation, convex-potential normalizing flows, optimal transport, and tran

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A Matched Spectral Benchmark of Quantum Inspired Feature Maps

DGX agent

arXiv:2605.24324v1 Announce Type: cross Abstract: Quantum machine learning is often motivated by the idea that quantum systems can expose useful high-dimensional structure that is difficult to access

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks

DGX agent

arXiv:2605.23977v1 Announce Type: new Abstract: This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS,

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

A Tabular Schedule Abstraction for Communication-Aware Evaluation of Pipeline-Parallel LLM Training

DGX agent

arXiv:2605.24006v1 Announce Type: cross Abstract: Pipeline parallelism is a key technique for distributed training of large language models because it reduces per-device parameter and activation memor

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

A tool to get Claude Code-style reliability from fully local models

DGX agent

Ollama exposes an Anthropic-compatible Messages endpoint , allowing developers to run powerful open-source AI models locally with no API costs and pair them with Claude Code for a capable local AI cod

model-releasesr-ollama
26 May 2026
Model Releases

A Two-Phase Stability Study of LLM Judges and Bar Council Examiners on Thai Bar-Exam Free-Form Essays

DGX agent

arXiv:2605.25652v1 Announce Type: new Abstract: Free-form legal essay evaluation in NLP treats expert inter-rater stability as a single ceiling number, and treats LLM-judge agreement with that ceiling

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

A World Model of Radiologist Reading for Medical Image Representation Learning

DGX agent

arXiv:2605.23992v1 Announce Type: cross Abstract: Radiologist eye-tracking data provide a rich record of how experts search, compare, and accumulate evidence during image reading; yet, existing method

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Acting on the Unseen: Communication-Free Collaborative Filtering for Decentralized Multi-Robot Task Allocation

DGX agent

arXiv:2605.25584v1 Announce Type: cross Abstract: Multi-robot task allocation usually assumes some combination of communication, known task models, or a coordinator. We study the opposite extreme, a r

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Action-Prior Denoising for Smooth Real-Time Chunking

DGX agent

arXiv:2605.25537v1 Announce Type: new Abstract: Real-time chunking (RTC) lets chunked action policies operate under inference delay by conditioning a newly generated action chunk on actions already co

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models

DGX agent

arXiv:2605.24011v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models exhibit remarkable action generation for embodied intelligence, but their heavy compute make deployment on edge pl

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Adaptive Graph Refinement and Label Propagation with LLMs for Cost-Effective Entity Resolution

DGX agent

arXiv:2605.25814v1 Announce Type: cross Abstract: Dirty entity resolution (ER), which identifies records referring to the same real-world entity from a single, messy dataset, is a fundamental task in

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Advancing Graph Few-Shot Learning via In-Context Learning

DGX agent

arXiv:2605.24410v1 Announce Type: new Abstract: Graph few-shot learning, which aims to classify nodes from novel classes with only a few labeled examples, is a widely studied problem in graph learning

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

DGX agent

arXiv:2605.23974v1 Announce Type: new Abstract: Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness its

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Agent World Model: Infinity Synthetic Environments for Agentic Reinforcement Learning

DGX agent

arXiv:2602.10090v3 Announce Type: replace Abstract: Recent advances in large language model (LLM) have empowered autonomous agents to perform multi-turn interactions with tools and environments. Howev

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

DGX agent

arXiv:2505.24876v2 Announce Type: replace-cross Abstract: Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understandi

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios

DGX agent

arXiv:2508.19988v3 Announce Type: replace Abstract: Large Language Models (LLMs) have achieved high accuracy on complex commonsense and mathematical problems that involve the composition of multiple r

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models

DGX agent

arXiv:2605.25901v1 Announce Type: cross Abstract: 3D Visual Grounding (3DVG) is an essential capability for embodied AI, requiring agents to localize objects in 3D scenes based on natural language des

model-releasesarxiv-cs-ro
26 May 2026
Model Releases

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

DGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

DGX agent

arXiv:2605.25358v1 Announce Type: cross Abstract: AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

DGX agent

arXiv:2605.25272v1 Announce Type: new Abstract: While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, ma

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AI Content Moderation in Therapy Conversations

DGX agent

arXiv:2605.25454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly being used for emotional support. They are also being developed for formal therapy purposes. However, LL

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications

DGX agent

arXiv:2602.22769v3 Announce Type: replace Abstract: Large Language Models (LLMs) are deployed as autonomous agents in increasingly complex applications, where enabling long-horizon memory is critical

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AME-TS: Anchored Mixture-of-Experts for Time Series Forecasting

DGX agent

arXiv:2605.25166v1 Announce Type: cross Abstract: Time series forecasting models are increasingly scaled through large Transformer backbones, yet most existing approaches process all series through a

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

An annoyance with Claude right now is that changes to the interface are badly documented, resulting in frustrating dead ends. For example, l…

DGX agent

An annoyance with Claude right now is that changes to the interface are badly documented, resulting in frustrating dead ends. For example, learning mode is migrating to a skill. Where is that skill? T

model-releasesethan-mollick--x
26 May 2026
← Previous
1…264265266267268…475
Next →