AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,628 results
Model Releases

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

DGX agent

arXiv:2606.23672v2 Announce Type: replace Abstract: This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In this tas

model-releasesarxiv-cs-ai
1 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer

DGX agent

arXiv:2606.31574v1 Announce Type: cross Abstract: Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service life of fusi

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

The automation rate of remote projects has increased ~4x in the past five months.

DGX agent

The automation rate of remote projects has increased ~4x in the past five months. New Remote Labor Index results: AI automation of real remote work is increasing fast. Claude Fable 5 now completes 16.

model-releasesdan-hendrycks--x
1 Jul 2026
Model Releases

The Bidirectional Process Reward Model

DGX agent

arXiv:2508.01682v3 Announce Type: replace Abstract: Process Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promi

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

The Calibration Turn in AI-Assisted Research: A Conceptual and Methodological Framework for Evidence-Licensed Claims

DGX agent

arXiv:2606.31273v1 Announce Type: new Abstract: AI-assisted research has entered a stage in which the central question is not only whether systems can generate hypotheses, run experiments, or produce

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

🔬 The Coolest Diffusion Research Isn't in LLMs — Evan Feinberg & Sergey Edunov, Genesis Molecular AI

DGX agent

This episode discusses cutting-edge diffusion model research applications beyond large language models, featuring insights from Evan Feinberg and Sergey Edunov of Genesis Molecular AI on how diffusion

model-releaseslatent-space
1 Jul 2026
Model Releases

The Decomposition Is the Fingerprint: Per-Component Identity for Agent Skills

DGX agent

arXiv:2606.31272v1 Announce Type: cross Abstract: AI agents increasingly acquire and execute skills at runtime: bundles of prompt instructions, executable code, and tool declarations fetched from mark

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

The Download: Anthropic launches Claude Science, and California’s carbon manure math

DGX agent

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Claude Science is Anthropic’s newest flagship product At an ev

model-releasesmit-tech-review
1 Jul 2026
Model Releases

The Great Algae Conspiracy Theory, sure to go down in history along with Hugo Chavez stole the 2020 election, the windmills are killing us, …

DGX agent

The Great Algae Conspiracy Theory, sure to go down in history along with Hugo Chavez stole the 2020 election, the windmills are killing us, and other Trump classics Trump on the Reflecting Pool: 'They

model-releasesanthropic--x
1 Jul 2026
Model Releases

The latest AI news we announced in June 2026

DGX agent

Google's June 2026 AI announcements included the launch of Gemini 3.5 Live Translate, the latest features in Android 17 and the new Google Home Speaker built for Gemini. The updates introduced new fea

model-releasesgoogle-ai
1 Jul 2026
Model Releases

Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action

DGX agent

arXiv:2606.31916v1 Announce Type: new Abstract: Theory of Mind (ToM) benchmarks for Large Language Models (LLMs) typically rely on passive question-answering formats, but the deployment of LLMs in inc

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Think in English, Answer in Korean: Efficient Adaptation of Multilingual Tool-Using Agents

DGX agent

arXiv:2606.31648v1 Announce Type: new Abstract: We present LuckyStar 111B, a 111B-parameter hybrid reasoning model developed through a collaboration between Cohere and LG CNS for Korean-English enterp

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

This is exactly why we believe in customization. Quick context: Factory's original secret scanner was deterministic, so it either flagged th…

DGX agent

This is exactly why we believe in customization. Quick context: Factory's original secret scanner was deterministic, so it either flagged things that weren't actually secrets (false positives) or miss

model-releasesfireworks-ai--x
1 Jul 2026
Model Releases

This is pretty concerning. You could still do this at the API level to some degree, but they seemingly just blatantly put it right into the …

DGX agent

This is pretty concerning. You could still do this at the API level to some degree, but they seemingly just blatantly put it right into the code? This is why open harnesses and agents are a much bette

model-releasesnous-research--x
1 Jul 2026
Model Releases

🚨 TRUMP’S FINANCIAL DISCLOSURE JUST DROPPED…AND IT’S WORSE THAN YOU THOUGHT. The U.S. Office of Government Ethics released Donald Trump’s 9…

DGX agent

🚨 TRUMP’S FINANCIAL DISCLOSURE JUST DROPPED…AND IT’S WORSE THAN YOU THOUGHT. The U.S. Office of Government Ethics released Donald Trump’s 927-page financial disclosure today. Here’s what a sitting U.S

model-releasesyann-lecun--x
1 Jul 2026
Model Releases

Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

DGX agent

arXiv:2606.31039v1 Announce Type: new Abstract: Large Language Models (LLMs) exhibit strong semantic capabilities, yet their resilience to manipulative linguistic patterns such as logical fallacies re

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

DGX agent

arXiv:2603.29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing ben

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Understanding and Evaluating Claw-like Agent Security Through a Computer-Systems Lens

DGX agent

arXiv:2606.30755v1 Announce Type: cross Abstract: Claw-like AI agents (e.g., OpenClaw) are always-on processes with persistent access to credentials, files, tools, and external services. They take on

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

UniCoder: Unified Visual-to-Code Generation via Symbolic Rewards and Reference-Guided Code Optimization

DGX agent

arXiv:2606.31732v1 Announce Type: new Abstract: Visual-to-Code generation, which transforms scientific plots, vector graphics, and webpages into executable scripts, demands a level of pixel-precise al

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

Unified Structural-Hydrodynamic Modeling of Underwater Underactuated Mechanisms and Soft Robots

DGX agent

arXiv:2603.07939v2 Announce Type: replace Abstract: Underwater robots are widely deployed for ocean exploration and manipulation. Underactuated mechanisms are advantageous in aquatic environments beca

model-releasesarxiv-cs-ro
1 Jul 2026
Model Releases

We are at a turning point: many of the decisions we make about AI today will permanently shape our future. Governments and the public need t…

DGX agent

We are at a turning point: many of the decisions we make about AI today will permanently shape our future. Governments and the public need to clearly understand the impacts, risks, and opportunities o

model-releasesyoshua-bengio--x
1 Jul 2026
Model Releases

We are SO back

DGX agent

We are SO back Claude Fable 5 will be available again globally tomorrow. After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to

model-releasesjerry-liu--x
1 Jul 2026
Model Releases

What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

DGX agent

arXiv:2606.31112v1 Announce Type: new Abstract: ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcripti

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

What If We Allocate Test-Time Compute Adaptively?

DGX agent

arXiv:2602.01070v5 Announce Type: replace Abstract: Test-time compute scaling allocates inference computation uniformly, uses fixed sampling strategies, and applies verification only for reranking. In

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

What Memory Do GUI Agents Really Need? From Passive Records to Active Task-Driving States

DGX agent

arXiv:2606.31612v1 Announce Type: new Abstract: Mobile GUI agents increasingly face long-horizon tasks that require reading, updating, and reusing task-relevant data across pages and applications. Exi

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models

DGX agent

arXiv:2606.30852v1 Announce Type: new Abstract: Reasoning models spend different amounts of useful computation across instances, but it remains unclear when a learned stopping rule improves over simpl

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

When LLMs Read Tables Carelessly: Measuring and Reducing Data Referencing Errors

DGX agent

arXiv:2606.32029v1 Announce Type: cross Abstract: While large language models (LLMs) perform well on table tasks, they still make data referencing errors (DREs), i.e., incorrectly citing or omitting t

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

When the Database Fails: Prompting LLM Dialogue Agents for Safe Recovery in Task-Oriented Dialogue

DGX agent

arXiv:2606.31307v1 Announce Type: new Abstract: Large language models used in task-oriented dialogue often produce fluent but unsafe responses when backend database calls fail, return empty results, o

model-releasesarxiv-cs-cl
1 Jul 2026
Model Releases

Why Solve It Twice? Hierarchical Accumulation of Skills for Transfer-Efficient ML Engineering

DGX agent

arXiv:2606.30911v1 Announce Type: new Abstract: ML engineering agents waste compute rediscovering known techniques because every competition is a cold start. We present HASTE, a hierarchical multi-age

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

DGX agent

arXiv:2606.31704v1 Announce Type: new Abstract: The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance dispari

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

Wiki!!!

DGX agent

Wiki!!! One unexpected outcome of this is that I'm now using the wiki as the ONLY place I run Claude Code I use it as a master controller for all of my repos, kicking off cross-repo tasks and using To

model-releasesharrison-chase--x
1 Jul 2026
Model Releases

Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models

DGX agent

arXiv:2606.31404v1 Announce Type: new Abstract: Human swarm intelligence demonstrates remarkable collective accuracy but faces scalability constraints in cost, coordination, and time. We investigate w

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Wordle 1,837 6/6 ⬛⬛⬛⬛⬛ ⬛⬛⬛⬛⬛ ⬛⬛🟨⬛⬛ ⬛🟩⬛⬛🟩 🟩🟩⬛⬛🟩 🟩🟩🟩🟩🟩

DGX agent

This entry documents a Wordle puzzle solution (puzzle #1,837) shared by Anthropic on X/Twitter, showing the complete sequence of guesses and letter feedback that led to solving the word on the sixth a

model-releasesanthropic--x
1 Jul 2026
Model Releases

Wordle 1,838 4/6 ⬛⬛⬛🟨🟨 ⬛⬛⬛⬛⬛ ⬛⬛🟨⬛🟩 🟩🟩🟩🟩🟩

DGX agent

This appears to be a Wordle game result shared by Anthropic on X (formerly Twitter), showing the solution was found in 4 attempts with a specific pattern of correct (green), present but misplaced (yel

model-releasesanthropic--x
1 Jul 2026
Model Releases

World-Model Collapse as a Phase Transition

DGX agent

arXiv:2606.31399v1 Announce Type: new Abstract: Water looks unchanged as it warms, then at a critical point it boils. We ask whether long-horizon language agents show an analogous transition in their

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

WorldRoamBench: An Open-World Benchmark for Long-Horizon Stability of Interactive World Models

DGX agent

arXiv:2606.31672v1 Announce Type: cross Abstract: Despite rapid progress in interactive world models (IWMs), existing benchmarks evaluate action following only at trajectory level and ignore memory an

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Xiaomi-GUI-0 Technical Report

DGX agent

arXiv:2606.31410v1 Announce Type: new Abstract: Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions s

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Yes! Pre-classifying routers are going to result in a lot of bad work because routing is hard and tend to underestimate the value of intelli…

DGX agent

Yes! Pre-classifying routers are going to result in a lot of bad work because routing is hard and tend to underestimate the value of intelligence on many problems. OpenAI learned this with GPT-5, now

model-releasesethan-mollick--x
1 Jul 2026
Model Releases

You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between…

DGX agent

You really need to benchmark models for your use case. As soon as judgements & decisions stack on top of each other, the differences between models amplifies, and no standard benchmark will tell you t

model-releasesethan-mollick--x
1 Jul 2026
Model Releases

Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models

DGX agent

arXiv:2606.31846v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models offer a promising framework for robotic manipulation by connecting language instructions, visual observations, and

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

3D Scene-Adaptive Trajectory-Controllable Human Image Animation with Camera Movement

DGX agent

arXiv:2606.30514v1 Announce Type: new Abstract: Human image animation, which aims to generate a video of a reference subject following a provided action sequence, has received increasing research inte

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

A Bayesian latent Gaussian process framework for aerodynamic uncertainty quantification

DGX agent

arXiv:2606.28871v1 Announce Type: cross Abstract: Predicting the aerodynamic performance (e.g. lift, drag, and moment coefficients) of an aircraft is challenging -- computational models are biased and

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

A Comparative Study on Affective Cues in Text Embeddings Across Psychological Emotion Theories

DGX agent

arXiv:2606.29068v1 Announce Type: cross Abstract: Text encoders are known for their utility in natural language processing, as they are able to efficiently compress inputs into dense vectors while pre

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

DGX agent

arXiv:2606.29719v1 Announce Type: cross Abstract: Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it.

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

A Machine-Verified Proof of a Quantum-Optimization Conjecture

DGX agent

arXiv:2606.29687v1 Announce Type: cross Abstract: We report a machine-verified resolution of a problem open for over a decade in quantum optimization: the Farhi, Goldstone and Gutmann (FGG) conjecture

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

A multi-architecture study of specificity refinement and false-positive mechanism analysis in prostate MRI

DGX agent

arXiv:2606.29977v1 Announce Type: cross Abstract: Objectives: To characterize residual false positives in prostate MRI detection, and to evaluate a lightweight post-hoc refinement head for case-level

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

A Multi-Dataset Benchmark for Evaluating LLM Agents in Microservice Failure Diagnosis

DGX agent

arXiv:2606.29193v1 Announce Type: cross Abstract: LLM-based agents are reshaping microservice operations into AgentOps, where benchmarks are key to evaluating failure diagnosis over multimodal observa

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models

DGX agent

arXiv:2606.28757v1 Announce Type: new Abstract: Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi

model-releasesarxiv-cs-cv
30 Jun 2026
← Previous
1…146147148149150…472
Next →