AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,561 results
Model Releases

The Oracle's Fingerprint: Correlated AI Forecasting Errors and the Limits of Bias Transmission

DGX agent

arXiv:2605.00844v1 Announce Type: cross Abstract: When large language models (LLMs) are consulted as forecasting tools, the independence of individual errors -- the foundation of collective intelligen

model-releasesarxiv-cs-ai
6 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development

DGX agent

arXiv:2605.01160v1 Announce Type: cross Abstract: Since 2022, AI-powered coding assistants have produced contradictory evidence: controlled studies report 20-56% productivity gains on well-scoped task

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It

DGX agent

arXiv:2605.03258v1 Announce Type: cross Abstract: Large language models often fail at simple counting tasks, even when the items to count are explicitly present in the prompt. We investigate whether t

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

DGX agent

arXiv:2605.03073v1 Announce Type: new Abstract: Niche-domain Indic ASR -- digit strings, currency amounts, addresses, brand names, English/Indic codemix -- is under-served by both open-source SOTA and

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

this deepagents deploy https://docs.langchain.com/oss/python/deepagents/deploy (or at least directionally where we want to take it) what's m…

DGX agent

this deepagents deploy https://docs.langchain.com/oss/python/deepagents/deploy (or at least directionally where we want to take it) what's missing? give us feedback! can someone PLEASE launch OS claud

model-releasesharrison-chase--x
6 May 2026
Model Releases

This week made something clear: you shouldn't take what most tech ceos are saying publicly seriously!

DGX agent

This week made something clear: you shouldn't take what most tech ceos are saying publicly seriously! From “Anthropic is Misanthropic” to “Claude is good for humanity and was impressed.” Most ironic o

model-releasesyann-lecun--x
6 May 2026
Model Releases

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

DGX agent

arXiv:2605.01809v1 Announce Type: cross Abstract: Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive medi

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Today we're releasing ZAYA1-8B, a reasoning MoE trained on @AMD and optimized for intelligence density. With <1B active params, it outperfor…

DGX agent

Today we're releasing ZAYA1-8B, a reasoning MoE trained on @AMD and optimized for intelligence density. With <1B active params, it outperforms open-weight models many times its size on math and reason

model-releasesclem-delangue--x
6 May 2026
Model Releases

Toward a Science of Intent: Closure Gaps and Delegation Envelopes for Open-World AI Agents

DGX agent

arXiv:2604.25000v2 Announce Type: replace Abstract: Recent work has framed intelligence in verifiable tasks as reducing time-to-solution through learned structure and test-time search, while systems w

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Toward Generative Quantum Utility via Correlation-Complexity Map

DGX agent

arXiv:2603.06440v2 Announce Type: replace Abstract: We study a practical question in generative quantum machine learning: given a classical dataset, can we determine, before training, whether it is we

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Towards Agentic Runtime Healing

DGX agent

arXiv:2408.01055v2 Announce Type: replace-cross Abstract: Self-healing systems have long been a focus of research, aiming to enable software to recover from unexpected runtime errors without human int

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Towards Multi-Agent Autonomous Reasoning in Hydrodynamics

DGX agent

arXiv:2605.01102v1 Announce Type: new Abstract: Single-agent systems (SAS) have become the default pattern for LLM-driven scientific workflows, but routing planning, tool use, and synthesis through a

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Towards Understanding Specification Gaming in Reasoning Models

DGX agent

arXiv:2605.02269v1 Announce Type: new Abstract: Specification gaming is a critical failure mode of LLM agents. Despite this, there has been little systematic research into when it arises and what driv

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Training-Free Probabilistic Time-Series Forecasting with Conformal Seasonal Pools

DGX agent

arXiv:2605.03789v1 Announce Type: cross Abstract: We propose Conformal Seasonal Pools (CSP), a training-free probabilistic time-series forecaster that mixes same-season empirical draws with signed res

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

DGX agent

arXiv:2605.03792v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar e

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration

DGX agent

arXiv:2605.01970v2 Announce Type: cross Abstract: Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characte

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Two frontier labs. One accelerated computing platform. Congrats to @SpaceX and @AnthropicAI on the new compute partnership, powered by 220,0…

DGX agent

Two frontier labs. One accelerated computing platform. Congrats to @SpaceX and @AnthropicAI on the new compute partnership, powered by 220,000+ NVIDIA GPUs inside Colossus 1. The future of AI runs on

model-releasesboris-cherny--x
6 May 2026
Model Releases

Two weeks after release, Hy3 preview is #1 on @OpenRouter's weekly leaderboard with 3.66T tokens processed, up 298% week-over-week. #1 in ov…

DGX agent

Two weeks after release, Hy3 preview is #1 on @OpenRouter's weekly leaderboard with 3.66T tokens processed, up 298% week-over-week. #1 in overall usage, tool calls, and coding. 15.4% market share acro

model-releasesjeremy-howard--x
6 May 2026
Model Releases

Uber uses OpenAI to help people earn smarter and book faster

DGX agent

Uber has integrated OpenAI's technology to enhance its platform, helping drivers optimize earnings through intelligent features and enabling users to book rides more efficiently. This partnership like

model-releasesopenai
6 May 2026
Model Releases

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning

DGX agent

arXiv:2605.03950v1 Announce Type: new Abstract: Although recent LMMs have become much stronger at visual perception, they remain unreliable on problems that require multi-step reasoning over visual ev

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Valley3: Scaling Omni Foundation Models for E-commerce

DGX agent

arXiv:2605.01278v1 Announce Type: new Abstract: In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understandi

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Vanishing L2 regularization for the softmax Multi Armed Bandit

DGX agent

arXiv:2605.03752v1 Announce Type: new Abstract: Multi Armed Bandit (MAB) algorithms are a cornerstone of reinforcement learning and have been studied both theoretically and numerically. One of the mos

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

@vasuman We've been saying this for months. The best compliment you can give an AI system is that it behaves exactly as expected. Every time…

DGX agent

@vasuman We've been saying this for months. The best compliment you can give an AI system is that it behaves exactly as expected. Every time. We built a campaign around it. It's called Boring AI: http

model-releasesai21-labs--x
6 May 2026
Model Releases

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

DGX agent

arXiv:2605.03276v1 Announce Type: new Abstract: Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage i

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

very fun to collab with @harvey on their Long Horizon Legal Agent Benchmark. We need more industry specific benchmarks, and Harvey is paving…

DGX agent

Harrison Chase expresses enthusiasm about collaborating with Harvey on their Long Horizon Legal Agent Benchmark, highlighting the value of developing industry-specific benchmarks for AI evaluation. Th

model-releasesharrison-chase--x
6 May 2026
Model Releases

💫Very happy to release NeuralBench, to benchmark Neuro AI models and datasets in the open! 🧵Thread, 💻Code, 📝White Paper below:

DGX agent

💫Very happy to release NeuralBench, to benchmark Neuro AI models and datasets in the open! 🧵Thread, 💻Code, 📝White Paper below: 🧠 Introducing NeuralBench: a unified, open-source framework to benchmark

model-releasesyann-lecun--x
6 May 2026
Model Releases

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

DGX agent

arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete '

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Vibe coding and agentic engineering are getting closer than I'd like

DGX agent

I recently talked with Joseph Ruscio about AI coding tools for Heavybit's High Leverage podcast: Ep. #9, The AI Coding Paradigm Shift with Simon Willison. Here are some of my highlights, including my

model-releasessimon-willison
6 May 2026
Model Releases

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models

DGX agent

arXiv:2605.01449v1 Announce Type: cross Abstract: Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, sug

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models

DGX agent

arXiv:2605.03351v1 Announce Type: new Abstract: Video vision-language models (VLMs) keep paying for visual state the stream already told us was stable. The factory wall did not move, but most VLM pipe

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

We also announced a compute partnership with SpaceX: 300+ MW of new capacity and 220K NVIDIA GPUs online within the month, all powering Clau…

DGX agent

Anthropic announced a major compute partnership with SpaceX involving over 300 MW of new computational capacity and 220,000 NVIDIA GPUs to be brought online within one month, dedicated to powering Cla

model-releasesboris-cherny--x
6 May 2026
Model Releases

we're continuing to see clear examples where a model's harness is a major determinant of overall performance. with the same model, running o…

DGX agent

we're continuing to see clear examples where a model's harness is a major determinant of overall performance. with the same model, running on same task, it's easy to observe very different scores depe

model-releasesharrison-chase--x
6 May 2026
Model Releases

We're winding back our peak hours limit reduction and doubling 5 hour limits. Excited to partner with SpaceX to bring you more compute and w…

DGX agent

We're winding back our peak hours limit reduction and doubling 5 hour limits. Excited to partner with SpaceX to bring you more compute and we'll keep pushing to bring you the best coding agent in the

model-releasesthariq--x
6 May 2026
Model Releases

We’ve agreed to a partnership with @SpaceX that will substantially increase our compute capacity. This, along with our other recent compute …

DGX agent

We’ve agreed to a partnership with @SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limi

model-releasesboris-cherny--x
6 May 2026
Model Releases

We’ve been working closely with the @harvey team on the launch of the Legal Agent Benchmark, a product focused on evaluating how open-weight…

DGX agent

We’ve been working closely with the @harvey team on the launch of the Legal Agent Benchmark, a product focused on evaluating how open-weight models perform on long-horizon, real-world legal tasks. Che

model-releasesfireworks-ai--x
6 May 2026
Model Releases

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-paramet…

DGX agent

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-parameter LLMs. With CuTeDSL integrated into our inference engine,

model-releasesperplexity--x
6 May 2026
Model Releases

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powe…

DGX agent

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powered by the LlamaParse TypeScript SDK 📱 Three steps, that’s i

model-releasesjerry-liu--x
6 May 2026
Model Releases

What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models

DGX agent

arXiv:2603.12799v2 Announce Type: replace Abstract: Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and chal

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

What's new in IAM: Security, governance, and runtime defense

DGX agent

The AI era demands a fundamental shift in security, and that includes identity and access management (IAM). Traditional controls simply aren’t built for autonomous AI agents that interact with sensiti

model-releasesgoogle-cloud-ai
6 May 2026
Model Releases

When Alignment Isn't Enough: Response-Path Attacks on LLM Agents

DGX agent

arXiv:2605.02187v1 Announce Type: cross Abstract: Bring-Your-Own-Key (BYOK) agent architectures let users route LLM traffic through third-party relays, creating a critical integrity gap: a malicious r

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

DGX agent

arXiv:2605.03096v1 Announce Type: cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

DGX agent

arXiv:2605.02915v1 Announce Type: new Abstract: Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its pr

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems

DGX agent

arXiv:2605.02463v1 Announce Type: cross Abstract: Multi-agent LLM systems are increasingly used to solve complex tasks through decomposition, debate, specialization, and ensemble reasoning. However, t

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Where to Bind Matters: Hebbian Fast Weights in Vision Transformers for Few-Shot Character Recognition

DGX agent

arXiv:2605.02920v1 Announce Type: cross Abstract: Standard transformer architectures learn fixed slow-weight representations during training and lack mechanisms for rapid adaptation within an episode.

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Wordle 1,781 3/6 ⬛🟨🟨⬛⬛ 🟩⬛⬛⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

Anthropic shared their Wordle game result for puzzle #1,781, solved in 3 attempts with a final correct answer shown by five green squares (all letters in correct positions). The emoji grid represents

model-releasesanthropic--x
6 May 2026
Model Releases

Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

DGX agent

arXiv:2605.03596v1 Announce Type: cross Abstract: Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

DGX agent

arXiv:2605.03475v1 Announce Type: new Abstract: Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal t

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

xAI and SpaceXAI have just made Colossus 1 available to Anthropic to support Claude. This means more than 220,000 NVIDIA GPUs in one of the …

DGX agent

xAI and SpaceXAI have just made Colossus 1 available to Anthropic to support Claude. This means more than 220,000 NVIDIA GPUs in one of the world’s largest and fastest-built AI superclusters are now h

model-releaseselon-musk--x
6 May 2026
← Previous
1…354355356357358…471
Next →