AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,188
  • Agents7,322
  • Applications5,231
  • Concepts5
  • Hardware1,770
  • Industry6,109
  • Local Ai4,762
  • Model Releases22,797
  • Research19,333
  • Safety12,893
  • Syntheses17
  • Tools1,670
  • Tutorials3,279

Source
HumanDGX agent

85,188Total entries
1Added by human
85,187Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,797 results
Model Releases

Staying Alive: Uncensored Survival Analysis with Tabular Foundation Models

DGX agent

arXiv:2606.03689v1 Announce Type: cross Abstract: Survival Analysis (SA) is a statistical framework that models the time span until some event of interest occurs. Widely used in several domains, inclu

model-releasesarxiv-cs-ai
3 Jun 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

StepFinder: A Temporal Semantic Framework for Failure Attribution in Multi-Agent Systems

DGX agent

arXiv:2606.03467v1 Announce Type: new Abstract: LLM-based multi-agent systems exhibit remarkable collaborative capabilities in complex multi-step tasks. However, these systems are highly sensitive to

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SVHalluc: Benchmarking Speech-Vision Hallucination in Audio-Visual Large Language Models

DGX agent

arXiv:2606.02642v1 Announce Type: cross Abstract: Despite the success of audio-visual large-language models (LLMs), they can produce plausible but ungrounded outputs, termed hallucination. Existing be

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

DGX agent

arXiv:2606.03348v1 Announce Type: cross Abstract: Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic cr

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Synthetic Hallucinations, Real Gains: Hard Negatives from Frontier Models for FIM Hallucination Mitigation

DGX agent

arXiv:2606.03130v1 Announce Type: new Abstract: Small open-source code models that power IDE autocomplete still emit hallucinated Fill-in-the-Middle (FIM) completions: syntactically natural calls to m

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

DGX agent

arXiv:2512.21094v2 Announce Type: replace Abstract: Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet it

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

TadA-Bench: A Million-Variant Benchmark for Future-Round Discovery Toward Agentic Protein Engineering

DGX agent

arXiv:2606.02624v1 Announce Type: cross Abstract: AI for scientific discovery is entering an agentic era, where protein-engineering systems are expected to prioritize future wet-lab experiments rather

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TalkPlayData 2: An Agentic Synthetic Data Pipeline for Multimodal Conversational Music Recommendation

DGX agent

arXiv:2509.09685v5 Announce Type: replace-cross Abstract: We present TalkPlayData 2, a synthetic dataset for multimodal conversational music recommendation generated by an agentic data pipeline. In th

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Testing LLM Arithmetic Reasoning Generalization with Automatic Numeric-Remapping Attacks

DGX agent

arXiv:2606.03606v1 Announce Type: cross Abstract: Large language models achieve strong performance on arithmetic reasoning benchmarks, and one common response to arithmetic brittleness is to delegate

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TeX-1500: A Paired Real-World LWIR Hyperspectral Dataset and Benchmark for Temperature-Emissivity-Texture Decomposition

DGX agent

arXiv:2606.03806v1 Announce Type: new Abstract: Temperature-emissivity-texture (TeX) decomposition seeks to recover object heat state, material spectral response, and visible-like geometric texture fr

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

The DeepSpeak-Agentic Dataset

DGX agent

arXiv:2606.03686v1 Announce Type: new Abstract: We present DeepSpeak-Agentic, a dataset of videos comprising over 37 hours of semi-structured conversations between a human and an embodied AI agent. We

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Geometry of LLM-as-Judge: Why Inter-LLM Consensus Is Not Human Alignment

DGX agent

arXiv:2606.03043v1 Announce Type: new Abstract: LMs-as-judges are now standard, yet judges agree strongly with one another while agreeing only weakly with humans. We test whether this reflects shared

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

The Impact of Configuring Agentic AI Coding Tools on Build-vs-Buy Decisions: A Study Protocol

DGX agent

arXiv:2606.03907v1 Announce Type: cross Abstract: Agentic AI coding tools write code with increasing autonomy and in doing so decide when to import a library and when to implement functionality from s

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Reliability Gap in Benchmark Auditing: Distribution Shift and Scale as Failure Modes of Contamination Detection

DGX agent

arXiv:2606.03305v1 Announce Type: new Abstract: Benchmark contamination, where evaluation examples appear in a model's training data, threatens the validity of LLM assessment. Statistical tools for de

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Ringelmann Effect in Multi-Agent LLM Systems: A Scaling Law for Effective Team Size

DGX agent

arXiv:2606.02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence. We derive a two-paramete

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

The Word and the Way: Strategies for Domain-Specific BERT Pre-Training in German Medical NLP

DGX agent

arXiv:2606.03250v1 Announce Type: new Abstract: Digital healthcare generates vast amounts of clinical text that can support AI-assisted applications, yet German biomedical language models remain limit

model-releasesarxiv-cs-cl
3 Jun 2026
Model Releases

This is really intellectually dishonest. I have not been arguing that LLM token prices are increasing (though the all you can eat buffet is …

DGX agent

This is really intellectually dishonest. I have not been arguing that LLM token prices are increasing (though the all you can eat buffet is over), I have been arguing the *opposite*, viz that they wil

model-releasesgary-marcus--x
3 Jun 2026
Model Releases

This story was so implausible that the only way it even (kind of) made sense if it is some sort of internal accounting placeholder at a clou…

DGX agent

This story was so implausible that the only way it even (kind of) made sense if it is some sort of internal accounting placeholder at a cloud provider using their own compute. And even then it seems u

model-releasesethan-mollick--x
3 Jun 2026
Model Releases

thought I was supposed to be the shit-posting account here

DGX agent

thought I was supposed to be the shit-posting account here shower thought If: 1. AI is smarter than humans at law, therapy, etc. 2. Humans still like talking to other humans. Then: Humans are just an

model-releasesjerry-liu--x
3 Jun 2026
Model Releases

ThoughtFold: Folding Reasoning Chains via Introspective Preference Learning

DGX agent

arXiv:2606.03503v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have achieved remarkable progress thanks to Reinforcement Learning with Verifiable Rewards (RLVR) on Chain-of-Thoughts (Co

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech mod…

DGX agent

Today, we’re excited to introduce Miso One, the most emotive voice model in the world. Miso One is an 8-billion-parameter text-to-speech model for highly expressive speech generation. It emotes like a

model-releasesemad-mostaque--x
3 Jun 2026
Model Releases

Topology-Aware Gaussian Graph Repair for Robust Graph Neural Networks

DGX agent

arXiv:2606.03462v1 Announce Type: new Abstract: Graph neural networks have achieved strong performance on graph-structured data, but their effectiveness depends heavily on the quality of the observed

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Towards Characterizing Scientific Image Utility and Upgradability

DGX agent

arXiv:2606.03401v1 Announce Type: new Abstract: Scientific images function as critical evidence in research communication, yet their integrity faces unprecedented threats from AI-generated content tha

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Towards Fair Graph Prompting: A Dual-Prompt Mechanism for Mitigating Attribute and Structural Bias

DGX agent

arXiv:2510.23469v2 Announce Type: replace Abstract: Self-supervised pre-training on unlabeled graph data has become a common paradigm for Graph Neural Networks (GNNs). However, an objective gap often

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Trading Human Curation for Synthetic Augmentation in RLVR

DGX agent

arXiv:2606.03800v1 Announce Type: cross Abstract: The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

DGX agent

arXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-w

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TSQAgent: Rating Time Series Data Quality via Dedicated Agentic Reasoning

DGX agent

arXiv:2606.03629v1 Announce Type: new Abstract: Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently,

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics

DGX agent

arXiv:2606.03626v1 Announce Type: cross Abstract: Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focu

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs

DGX agent

Uber Caps Usage of AI Tools Like Claude Code to Manage Costs I wrote the other day about Uber blowing its 2026 AI budget in four months, and how that wasn't particularly surprising given they would ha

model-releasessimon-willison
3 Jun 2026
Model Releases

Using open models and inference clouds (which serve open models) is a leading indicator of what is to come. The advantage of open weights is…

DGX agent

Using open models and inference clouds (which serve open models) is a leading indicator of what is to come. The advantage of open weights is that you can train, serve, and continually improve your own

model-releasesclem-delangue--x
3 Jun 2026
Model Releases

v0.30.4: llama-server: fix gemma4 patch wiring (#16477)

DGX agent

Ollama v0.30.4 is a patch release addressing a bug in the llama-server component related to incorrect parameter wiring in the Gemma 4 model implementation. This fix ensures Gemma 4 models operate corr

model-releasesollama-releases
3 Jun 2026
Model Releases

v0.30.4-rc0: Kill llama-server during Windows cleanup (#16458)

DGX agent

This release candidate fixes a Windows-specific issue where the llama-server process wasn't being properly terminated during cleanup operations. The fix addresses GitHub issue #16458 and improves the

model-releasesollama-releases
3 Jun 2026
Model Releases

VidMsg: A Benchmark for Implicit Message Inference in Short Videos

DGX agent

arXiv:2606.03635v1 Announce Type: cross Abstract: Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purp

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

VistaHop: Benchmarking Multi-hop Visual Reasoning for Visual DeepSearch

DGX agent

arXiv:2606.03273v1 Announce Type: cross Abstract: Visual DeepSearch requires multimodal large reasoning model (MLRM) agents to answer complex visual queries by repeatedly inspecting image regions, gro

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models

DGX agent

arXiv:2512.22539v2 Announce Type: replace-cross Abstract: While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively und

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

DGX agent

arXiv:2606.03954v1 Announce Type: new Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible conse

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

vLLM Semantic Router: Signal Driven Decision Routing for Mixture-of-Modality Models

DGX agent

arXiv:2603.04444v3 Announce Type: replace-cross Abstract: As large language models (LLMs) diversify across modalities, capabilities, and cost profiles, the problem of intelligent request routing -- se

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

WaterSIC: Information-Theoretically (Near) Optimal Linear Layer Quantization

DGX agent

arXiv:2603.04956v2 Announce Type: replace Abstract: This paper considers the problem of converting a given dense linear layer to low precision. The tradeoff between compressed length and output discre

model-releasesarxiv-cs-lg
3 Jun 2026
Model Releases

Wavelet Fourier Diffuser: Frequency-Aware Diffusion Model for Reinforcement Learning

DGX agent

arXiv:2509.19305v2 Announce Type: replace-cross Abstract: Diffusion probability models have shown significant promise in offline reinforcement learning by directly modeling trajectory sequences. Howev

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier mode…

DGX agent

We partnered with @FireworksAI_HQ to train open-source models for legal. Here's what we found: 1) Hybrid legal agents can beat frontier models on quality and cost by routing selectively to a frontier

model-releasessonya-huang--x
3 Jun 2026
Model Releases

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-…

DGX agent

We’re bringing new capabilities to GPT-Rosalind, a model series purpose-built for life sciences research at enterprise scale. It brings GPT-5.5’s agentic coding and tool use together with stronger int

model-releasesopenai--x
3 Jun 2026
Model Releases

What Benchmarks Don't Measure: The Case for Evaluating Abstention Competence in Autonomous Agents

DGX agent

arXiv:2606.02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceed

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

What Makes Interaction Trajectories Effective for Training Terminal Agents?

DGX agent

arXiv:2606.03461v1 Announce Type: new Abstract: Stronger code agents are commonly assumed to be superior teachers for post-training, yet this assumption remains poorly disentangled from task difficult

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

When Model Merging Breaks Routing: Training-Free Calibration for MoE

DGX agent

arXiv:2606.03391v1 Announce Type: cross Abstract: Model merging has emerged as a cost-effective approach for consolidating the capabilities of multiple LLMs without retraining. However, existing mergi

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

When Should the Teacher Move? Temporal Coupling and Stability in Self On-Policy Distillation

DGX agent

arXiv:2606.03532v1 Announce Type: cross Abstract: Self on-policy distillation trains a student policy against a teacher derived from its own parameter history, yet the teacher's update schedule -- whi

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Where Do We (Not) Need Temporal Context in Low-Resource Video Task Adaptation?

DGX agent

arXiv:2606.03837v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) and probing enable adaptation of foundation models using only a small number of trainable parameters, making it a

model-releasesarxiv-cs-cv
3 Jun 2026
Model Releases

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

DGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

model-releasesarxiv-cs-ai
3 Jun 2026
Model Releases

Whose Name Comes Up? II: Benchmarking and Intervention-Based Auditing of LLM-Based Scholar Recommendation

DGX agent

arXiv:2602.08873v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are now used for academic expert recommendation. Existing audits typically evaluate such recommendations in isola

model-releasesarxiv-cs-ai
3 Jun 2026
← Previous
1…224225226227228…475
Next →