AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,115
  • Agents7,313
  • Applications5,228
  • Concepts5
  • Hardware1,762
  • Industry6,105
  • Local Ai4,756
  • Model Releases22,759
  • Research19,333
  • Safety12,889
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
85,115Total entries
1Added by human
85,114Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,767 results
Model Releases

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages

DGX agent

arXiv:2605.27901v1 Announce Type: cross Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. Howeve

model-releasesarxiv-cs-ai
28 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

DGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

The Importance of Being Statistically Earnest: A Critical Re-evaluation of GSM-Symbolic

DGX agent

arXiv:2605.28700v1 Announce Type: new Abstract: The GSM-Symbolic benchmark (Mirzadeh et al., 2025) reported consistent performance drops across 25 Large Language Models (LLMs) when tested on template-

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

DGX agent

arXiv:2605.28020v1 Announce Type: new Abstract: With the rapid progress of large language models (LLMs), reliably evaluating the capabilities of pre-trained LLMs has become increasingly important. The

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

The Name’s Gaming … Cloud Gaming: ‘007 First Light’ Launches on GeForce NOW

DGX agent

License to stream, shaken and stirred. GeForce NOW is dialing up the espionage with the launch of 007 First Light, letting members slip into James Bond’s reimagined origin story from almost any device

model-releasesnvidia-blog
28 May 2026
Model Releases

The Point, the Vision and the Text: Does Point Cloud Boost Spatial Reasoning of Large Language Models? A Bias-Controlled Study

DGX agent

arXiv:2504.04540v2 Announce Type: replace-cross Abstract: 3D Large Language Models (LLMs) leveraging spatial information in point clouds for 3D spatial reasoning attract great attention. Despite some

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

The Script is All You Need: An Agentic Framework for Long-Horizon Dialogue-to-Cinematic Video Generation

DGX agent

arXiv:2601.17737v3 Announce Type: replace-cross Abstract: Recent advances in video generation have produced models capable of synthesizing stunning visual content from simple text prompts. However, th

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Thermodynamic properties of chemically disordered compounds via AI-driven estimation of partition function with the PULSE method

DGX agent

arXiv:2605.28594v1 Announce Type: cross Abstract: In this article, we present an improved version of the PULSE method (Partition function Unsupervised Learning Sampling and Evaluation) for estimating

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

this is so funny, training opus 4.7 on business skills makes it misaligned and dishonest 😭

DGX agent

this is so funny, training opus 4.7 on business skills makes it misaligned and dishonest 😭 Learnings from testing Claude Opus 4.8: > Much worse than Opus 4.7 and GPT 5.5 on Vending Bench > More aligne

model-releasesemad-mostaque--x
28 May 2026
Model Releases

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to A…

DGX agent

Too many business leaders believe that AI says what it means. And it’s odd because we naturally attribute a high number of human traits to AI, and yet we refuse to believe it can have hidden intent? 3

model-releasesallie-k--miller--x
28 May 2026
Model Releases

Took some inspiration from @vboykis and converted my first ever talk into a blog post. I talk about the role of agentic search in context en…

DGX agent

Took some inspiration from @vboykis and converted my first ever talk into a blog post. I talk about the role of agentic search in context engineering. Together we build an intuition on the strengths a

model-releasesswyx--x
28 May 2026
Model Releases

Tool Forge: A Validation-Carrying Toolchain for Governed Agentic Execution

DGX agent

arXiv:2605.28000v1 Announce Type: cross Abstract: Large language model agents are increasingly expected to perform operational work: calling APIs, manipulating files, assembling workflows, and acting

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Towards Faithful Agentic XAI: A Verification Method and an Open-World Benchmark for Better Model Faithfulness

DGX agent

arXiv:2605.27879v1 Announce Type: new Abstract: Explainable AI (XAI) helps users interpret model behavior and identify potential faults. Agentic XAI systems use Large Language Models (LLMs) to make ex

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

TRACER: Turn-level Regret Matching with Inner Reinforcement Credit for Cooperative Multi-LLM Reasoning

DGX agent

arXiv:2605.28699v1 Announce Type: new Abstract: Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain d

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Trinity: Unifying Class-Agnostic Terrain and Semantic Segmentation for Unstructured Outdoor Environments by Leveraging Synthetic Data

DGX agent

arXiv:2605.27644v1 Announce Type: cross Abstract: Terrain understanding is fundamental for mobile robots operating in unstructured outdoor environments. Existing vision-based traversability estimation

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Understanding Generalization and Forgetting in In-Context Continual Learning

DGX agent

arXiv:2605.28705v1 Announce Type: new Abstract: In-context learning (ICL) derives its power from enabling Large Language Models to adapt to new tasks via prompt-based reasoning alone, entirely bypassi

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

DGX agent

arXiv:2605.28063v1 Announce Type: cross Abstract: Audio generation has made significant progress, yet synthesizing unified audio where speech and sounds are naturally composited remains a challenge. C

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

UniMaia: Steering Chess Policies with Language for Human-like Play

DGX agent

arXiv:2605.27767v1 Announce Type: cross Abstract: Recent advances in large language models have enabled natural language to serve as a flexible interface for controlling complex systems, but often at

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis

DGX agent

arXiv:2605.27401v1 Announce Type: cross Abstract: There is a growing interest in utilizing synthetic populations for a diverse range of applications. At the same time, we are witnessing a tremendous g

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Verifiable Benchmarking of Long-Horizon Spatial Biology

DGX agent

arXiv:2605.28065v1 Announce Type: new Abstract: AI agents are increasingly useful for biological data analysis, but existing benchmarks mostly test broad biological knowledge, executable workflows, or

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VeriTrip: A Verifiable Benchmark for Travel Planning Agents over Unstructured Web Corpora

DGX agent

arXiv:2605.28683v1 Announce Type: new Abstract: Existing benchmarks have laid the foundation for travel planning agents by establishing API-centric paradigms. However, as the capabilities of Autonomou

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild

DGX agent

arXiv:2605.27882v1 Announce Type: cross Abstract: LLM-based agents score well on search benchmarks, yet real users consistently find results unsatisfying, revealing a persistent evaluation-experience

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning

DGX agent

arXiv:2510.08555v2 Announce Type: replace Abstract: Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpaint

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs

DGX agent

arXiv:2605.28422v1 Announce Type: cross Abstract: Latent reasoning enables reasoning over continuous hidden states rather than explicit tokens, avoiding the language bottleneck and inference overhead

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

VULPO: Context-Aware Vulnerability Detection via On-Policy LLM Optimization

DGX agent

arXiv:2511.11896v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have recently shown strong potential in vulnerability detection (VD). However, accurately detecting vulnerabiliti

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

wait… if most people think 5.5 is better than 4.7, i assume that’s due to terminal coding benchmark… 4.8 is still outperformed by 5.5

DGX agent

wait… if most people think 5.5 is better than 4.7, i assume that’s due to terminal coding benchmark… 4.8 is still outperformed by 5.5 Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper ju

model-releasesjeremy-howard--x
28 May 2026
Model Releases

We also shipped dynamic workflows in Claude Code (research preview), for tasks too big for one pass. Make sure to default to auto mode so Cl…

DGX agent

We also shipped dynamic workflows in Claude Code (research preview), for tasks too big for one pass. Make sure to default to auto mode so Claude isn't stopping for permissions. It's token-intensive, s

model-releasesboris-cherny--x
28 May 2026
Model Releases

Weak Convergence Analysis of Online Neural Actor-Critic Algorithms

DGX agent

arXiv:2403.16825v2 Announce Type: replace Abstract: We prove that a single-layer neural network trained with the online actor critic algorithm converges in distribution to a random ordinary differenti

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

WeatherCity: Urban Scene Reconstruction with Controllable Multi-Weather Transformation

DGX agent

arXiv:2602.22096v2 Announce Type: replace Abstract: Editable high-fidelity 4D scenes are crucial for autonomous driving, as they can be applied to end-to-end training and closed-loop simulation. Howev

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

We're releasing Paris 2.0, which, to our knowledge, is the world's first decentralized trained video generation model. We benchmarked it aga…

DGX agent

We're releasing Paris 2.0, which, to our knowledge, is the world's first decentralized trained video generation model. We benchmarked it against a monolithic model trained on the same data and compute

model-releasesclem-delangue--x
28 May 2026
Model Releases

We've raised 65 billion in Series H funding at a 965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and @sequo…

DGX agent

We've raised 65 billion in Series H funding at a 965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and @sequoia. This investment will help us advance our research and expa

model-releasessonya-huang--x
28 May 2026
Model Releases

What-If World: A Causal Benchmark for General World Models in Embodied Scenarios

DGX agent

arXiv:2605.27589v1 Announce Type: new Abstract: Video generation models are increasingly used as world simulators for tasks like driving and robotic manipulation. What matters in these settings is not

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

When Context Flips, Safety Breaks: Diagnosing Brittle Safety in Aligned Language Models

DGX agent

arXiv:2605.27851v1 Announce Type: new Abstract: Safety benchmark scores provide incomplete evidence of deployment readiness: aligned language models often adhere to rigid rules even when a situational

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

When do complex-valued neural networks help? A study of representation, geometry, and optimization

DGX agent

arXiv:2605.27673v1 Announce Type: new Abstract: Complex-valued Neural Networks (CVNNs) are often motivated by domains where information is naturally encoded in magnitude and phase. Yet complex-valued

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

When Interpretability Is Unequally Distributed: Fairness in Hybrid Interpretable Models

DGX agent

arXiv:2605.28626v1 Announce Type: new Abstract: Hybrid interpretable models combine a transparent component with a black-box model by assigning some examples to the former and deferring the rest to th

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Who's going first

DGX agent

Who's going first Introducing Claude Opus 4.8: it builds on Opus 4.7 with sharper judgment, more honesty about its own progress, and the ability to work independently for longer than its predecessors.

model-releasesjerry-liu--x
28 May 2026
Model Releases

Whose Name Comes Up? III: Persona Prompting Effects in LLM-Based Scholar Recommendation

DGX agent

arXiv:2605.28187v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as scholar recommenders, shaping who is seen as an expert in academia. Existing audits remain Engli

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Why LLMs Fail at Causal Discovery and How Interventional Agents Escape

DGX agent

arXiv:2605.27567v1 Announce Type: new Abstract: Causal discovery is a cornerstone of scientific reasoning, yet whether large language models can perform it reliably remains an open question. Recent be

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Wordle 1,803 5/6 🟩🟩⬛⬛⬛ ⬛⬛⬛⬛⬛ 🟩🟩🟩⬛⬛ 🟩🟩🟩⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

This entry documents a Wordle game result (puzzle #1,803) where the player achieved a solution in 5 out of 6 allowed guesses. The color-coded emoji sequence shows the player's guess progression, with

model-releasesanthropic--x
28 May 2026
Model Releases

XTransfer: Modality-Agnostic Few-Shot Model Transfer for Human Sensing at the Edge

DGX agent

arXiv:2506.22726v4 Announce Type: replace Abstract: Deep learning for human sensing on edge systems presents significant potential for smart applications. However, its training and development are hin

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

You Live More Than Once: Towards Hierarchical Skill Meta-Evolving

DGX agent

arXiv:2605.28390v1 Announce Type: new Abstract: Test-time skill evolving is regarded as a new paradigm for enhancing deployed agentic systems. Existing works mainly focus on hard-coded skill evolving

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

You Only Align Once: Propagating Cooperative Behaviors in Multi-Agent Systems through Seed Agents

DGX agent

arXiv:2605.27586v1 Announce Type: cross Abstract: Ensuring agent behaviors in distributed open multi-agent systems remains challenging, especially as populations grow and unaligned agents may exist. W

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

ZipRL: Adaptive Multi-Turn Context Compression with Hindsight Response Replay

DGX agent

arXiv:2605.28069v1 Announce Type: new Abstract: Adaptive context compression is vital for scaling Large Language Models (LLMs) to complex, multi-turn agent tasks. However, rule-based compression metho

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

A closer look at what we released today 🧵 ESMC is a language model trained on billions of protein sequences spanning the full diversity of …

DGX agent

A closer look at what we released today 🧵 ESMC is a language model trained on billions of protein sequences spanning the full diversity of life. Trained across 2.8 billion sequences, the model is expo

model-releasesyann-lecun--x
27 May 2026
Model Releases

A Dataset of Robot-Patient and Doctor-Patient Medical Dialogues for Spoken Language Processing Tasks

DGX agent

arXiv:2605.26747v1 Announce Type: new Abstract: Large Language Models (LLMs) have brought huge improvements to Artificial Intelligence (AI), which can be applied to general-purpose tasks. However, the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing (Steven Levy/Wired)

DGX agent

Steven Levy / Wired: A deep dive into how Anthropic's Claude Code and Peter Steinberger's OpenClaw unleashed the AI agent revolution that is rapidly transforming modern computing — The definitive stor

model-releasestechmeme
27 May 2026
Model Releases

A Deep State-Space Model Compression Method using Upper Bound on Output Error

DGX agent

arXiv:2510.14542v2 Announce Type: replace-cross Abstract: We study deep state-space models (Deep SSMs) that contain linear quadratic-output (LQO) systems as internal blocks and present a compression m

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

A Guide to AI Cold Starts on Cloud Run

DGX agent

I saw a developer asking on Reddit if there was any “sane way” to manage Cloud Run cold starts for AI across multiple regions. They were experiencing startup latencies of up to 20 seconds, a frustrati

model-releasesgoogle-cloud-ai
27 May 2026
← Previous
1…258259260261262…475
Next →