AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
All
85,115Total entries
1Added by human
85,114Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,759 results
Model Releases

this is so funny, training opus 4.7 on business skills makes it misaligned and dishonest 😭

DGX agent

this is so funny, training opus 4.7 on business skills makes it misaligned and dishonest 😭 Learnings from testing Claude Opus 4.8: > Much worse than Opus 4.7 and GPT 5.5 on Vending Bench > More aligne

model-releasesemad-mostaque--x
28 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to A…

DGX agent

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to AI, and yet we refuse to believe it can have hidden intent? 3

model-releasesallie-k--miller--x
28 May 2026
Model Releases

Took some inspiration from @vboykis and converted my first ever talk into a blog post. I talk about the role of agentic search in context en…

DGX agent

Took some inspiration from @vboykis and converted my first ever talk into a blog post. I talk about the role of agentic search in context engineering. Together we build an intuition on the strengths a

model-releasesswyx--x
28 May 2026
Model Releases

Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution

DGX agent

arXiv:2605.28000v1 Announce Type: cross Abstract: Large language model agents are increasingly expected to perform operational work: calling APIs, manipulating files, assembling workflows, and acting

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

DGX agent

arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make ex

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning

DGX agent

arXiv:2605.28699v1 Announce Type: new Abstract: Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain d

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Trinity: Unifying Class-Agnostic Terrain and Semantic Segmentation for Unstructured Outdoor Environments by Leveraging Synthetic Data

DGX agent

arXiv:2605.27644v1 Announce Type: cross Abstract: Terrain understanding is fundamental for mobile robots operating in unstructured outdoor environments. Existing vision-based traversability estimation

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Understanding Generalization and Forgetting in In-Context Continual Learning

DGX agent

arXiv:2605.28705v1 Announce Type: new Abstract: In-context learning (ICL) derives its power from enabling Large Language Models to adapt to new tasks via prompt-based reasoning alone, entirely bypassi

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

DGX agent

arXiv:2605.28063v1 Announce Type: cross Abstract: Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. C

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

UniMaia: Steering Chess Policies with Language for Human-like Play

DGX agent

arXiv:2605.27767v1 Announce Type: cross Abstract: Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis

DGX agent

arXiv:2605.27401v1 Announce Type: cross Abstract: There is a growing interest in utilizing synthetic populations for a diverse range of applications. At the same time, we are witnessing a tremendous g

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Verifiable Benchmarking of Long-Horizon Spatial Biology

DGX agent

arXiv:2605.28065v1 Announce Type: new Abstract: AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

DGX agent

arXiv:2605.28683v1 Announce Type: new Abstract: Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomou

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

DGX agent

arXiv:2605.27882v1 Announce Type: cross Abstract: LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

DGX agent

arXiv:2510.08555v2 Announce Type: replace Abstract: Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpaint

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

DGX agent

arXiv:2605.28422v1 Announce Type: cross Abstract: Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization

DGX agent

arXiv:2511.11896v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently shown strong potential in vulnerability detection (VD). However, accurately detecting vulnerabiliti

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

wait… if most people think 5.5 is better than 4.7, i assume that’s due to terminal coding benchmark… 4.8 is still outperformed by 5.5

DGX agent

wait… if most people think 5.5 is better than 4.7, i assume that’s due to terminal coding benchmark… 4.8 is still outperformed by 5.5 Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper ju

model-releasesjeremy-howard--x
28 May 2026
Model Releases

We also shipped dynamic workflows in Claude Code (research preview), for tasks too big for one pass. Make sure to default to auto mode so Cl…

DGX agent

We also shipped dynamic workflows in Claude Code (research preview), for tasks too big for one pass. Make sure to default to auto mode so Claude isn't stopping for permissions. It's token-intensive, s

model-releasesboris-cherny--x
28 May 2026
Model Releases

Weak Convergence Analysis of Online Neural Actor-Critic Algorithms

DGX agent

arXiv:2403.16825v2 Announce Type: replace Abstract: We prove that a single-layer neural network trained with the online actor critic algorithm converges in distribution to a random ordinary differenti

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

DGX agent

arXiv:2602.22096v2 Announce Type: replace Abstract: Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. Howev

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

We're releasing Paris 2.0, which, to our knowledge, is the world's first decentralized trained video generation model. We benchmarked it aga…

DGX agent

We're releasing Paris 2.0, which, to our knowledge, is the world's first decentralized trained video generation model. We benchmarked it against a monolithic model trained on the same data and compute

model-releasesclem-delangue--x
28 May 2026
Model Releases

We've raised 65 billion in Series H funding at a 965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and @sequo…

DGX agent

We've raised 65 billion in Series H funding at a 965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and @sequoia. This investment will help us advance our research and expa

model-releasessonya-huang--x
28 May 2026
Model Releases

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios

DGX agent

arXiv:2605.27589v1 Announce Type: new Abstract: Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

DGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When do complex-valued neural networks help? A study of representation, geometry, and optimization

DGX agent

arXiv:2605.27673v1 Announce Type: new Abstract: Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

DGX agent

arXiv:2605.28626v1 Announce Type: new Abstract: Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to th

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Who's going first

DGX agent

Who's going first Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors.

model-releasesjerry-liu--x
28 May 2026
Model Releases

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

DGX agent

arXiv:2605.28187v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as scholar recommenders, shaping who is seen as an expert in academia. Existing audits remain Engli

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Why LLMs Fail at Causal Discovery and How Interventional Agents Escape

DGX agent

arXiv:2605.27567v1 Announce Type: new Abstract: Causal discovery is a cornerstone of scientific reasoning, yet whether large language models can perform it reliably remains an open question. Recent be

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Wordle 1,803 5/6 🟩🟩⬛⬛⬛ ⬛⬛⬛⬛⬛ 🟩🟩🟩⬛⬛ 🟩🟩🟩⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

This entry documents a Wordle game result (puzzle #1,803) where the player achieved a solution in 5 out of 6 allowed guesses. The color-coded emoji sequence shows the player's guess progression, with

model-releasesanthropic--x
28 May 2026
Model Releases

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

DGX agent

arXiv:2506.22726v4 Announce Type: replace Abstract: Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hin

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

DGX agent

arXiv:2605.28390v1 Announce Type: new Abstract: Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

DGX agent

arXiv:2605.27586v1 Announce Type: cross Abstract: Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. W

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

DGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

A closer look at what we released today 🧵 ESMC is a language model trained on billions of protein sequences spanning the full diversity of …

DGX agent

A closer look at what we released today 🧵 ESMC is a language model trained on billions of protein sequences spanning the full diversity of life. Trained across 2.8 billion sequences, the model is expo

model-releasesyann-lecun--x
27 May 2026
Model Releases

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

DGX agent

arXiv:2605.26747v1 Announce Type: new Abstract: Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing (Steven Levy/Wired)

DGX agent

Steven Levy / Wired: A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing — The definitive stor

model-releasestechmeme
27 May 2026
Model Releases

A Deep State-Space Model Compression Method using Upper Bound on Output Error

DGX agent

arXiv:2510.14542v2 Announce Type: replace-cross Abstract: We study deep state-space models (Deep SSMs) that contain linear quadratic-output (LQO) systems as internal blocks and present a compression m

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

A Guide to AI Cold Starts on Cloud Run

DGX agent

I saw a developer asking on Reddit if there was any “sane way” to manage Cloud Run cold starts for AI across multiple regions. They were experiencing startup latencies of up to 20 seconds, a frustrati

model-releasesgoogle-cloud-ai
27 May 2026
Model Releases

A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection

DGX agent

arXiv:2605.26533v1 Announce Type: cross Abstract: Automated industrial inspection requires both precise defect localization and structured maintenance report generation; in current practice these task

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

A newly released AI tool has generated an atlas of more than one billion predicted protein structures and billions more protein sequences. h…

DGX agent

DeepMind's AlphaFold3 and related tools have generated a comprehensive atlas containing over one billion predicted protein structures and additional billions of protein sequences, representing a major

model-releasesyann-lecun--x
27 May 2026
Model Releases

AdaSD: Adaptive Speculative Decoding for Efficient Language Model Inference

DGX agent

arXiv:2512.11280v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved remarkable performance across a wide range of tasks, but their increasing parameter sizes significantly s

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

ADRD-Bench: A Preliminary LLM Benchmark for Alzheimer's Disease and Related Dementias

DGX agent

arXiv:2602.11460v2 Announce Type: replace Abstract: Large language models (LLMs) have shown great potential for healthcare applications. However, existing evaluation benchmarks provide minimal coverag

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Advancing Creative Physical Intelligence in Large Multimodal Models

DGX agent

arXiv:2605.26396v1 Announce Type: new Abstract: Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to d

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

AgentSociety: Incentivizing Agentic Social Intelligence

DGX agent

arXiv:2605.26203v1 Announce Type: cross Abstract: The success of deployed agents relies on their ability to handle open-ended user requests using their inherent capabilities, not only in solving reque

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

AGORA: Adapter-Grounded Observation-Action Retention for Inference-Free Prompt Compression in LLM Agents

DGX agent

arXiv:2605.26596v1 Announce Type: new Abstract: The token-level extractive compressors widely used for general LM context are structurally inappropriate for LLM agents: across 17 (env, backbone, metho

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

AI evaluation may bias perceptions: The importance of context in interpreting academic writing

DGX agent

arXiv:2605.26662v1 Announce Type: cross Abstract: This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries

model-releasesarxiv-cs-ai
27 May 2026
← Previous
1…257258259260261…475
Next →