AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,548
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,751
  • Industry6,096
  • Local Ai4,728
  • Model Releases22,555
  • Research19,193
  • Safety12,813
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,548Total entries
1Added by human
84,547Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,563 results
8 Jul 2026

ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations

Model ReleasesDGX agent

arXiv:2607.06052v1 Announce Type: new Abstract: Humanoid robots are increasingly expected to perform contact-rich tasks that require not only accurate whole-body motion but also robust physical intera

To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software eng…

Model ReleasesDGX agent

To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software engineers. That helped us examine tasks at scale while keeping

Tomorrow at 10am PT I'm hosting a live walkthrough of how we progressed from single-player Claude Code to multi-player Claude Tag. Then, we'…


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

Tomorrow at 10am PT I'm hosting a live walkthrough of how we progressed from single-player Claude Code to multi-player Claude Tag. Then, we're going deep on how Claude Tag actually works. AI used to f

Trained entirely in simulation: ~400,000 trajectories across 6,000 scenes. A prefix-caching recipe cuts training tokens by 22×, turning mont…

Model ReleasesDGX agent

Trained entirely in simulation: ~400,000 trajectories across 6,000 scenes. A prefix-caching recipe cuts training tokens by 22×, turning months-long runs into days. Online RL (CISPO) pushes success rat

Transformers converge to invariant algorithmic cores

Model ReleasesDGX agent

arXiv:2602.22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function. Studying any single trained neural n

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

Model ReleasesDGX agent

arXiv:2607.06306v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex pro

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation

Model ReleasesDGX agent

arXiv:2607.06537v1 Announce Type: new Abstract: Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity

VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection

Model ReleasesDGX agent

arXiv:2607.06254v1 Announce Type: cross Abstract: Deepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models (LLMs), and

Verification of Dynamic Holographic Behavior in Identity Documents

Model ReleasesDGX agent

arXiv:2607.06466v1 Announce Type: new Abstract: This paper addresses the remote verification of the authenticity of Optically Variable Devices (commonly known as holograms) on identity documents. Typi

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

Model ReleasesDGX agent

arXiv:2510.13808v2 Announce Type: replace Abstract: Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel doma

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capabil…

Model ReleasesDGX agent

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and ar

We need AI model selection to be MUCH easier ASAP. I want proactive flags from my AI systems suggesting models. I want my AI harness to say …

Model ReleasesDGX agent

We need AI model selection to be MUCH easier ASAP. I want proactive flags from my AI systems suggesting models. I want my AI harness to say 'hey allie, my girl, you keep asking for bar recommendations

We tuned the harness for @NVIDIAAI Nemotron 3 Ultra. Benchmark-leading performance. 10x lower inference costs. ✅ An aggregate score of 0.86 …

Model ReleasesDGX agent

We tuned the harness for @NVIDIAAI Nemotron 3 Ultra. Benchmark-leading performance. 10x lower inference costs. ✅ An aggregate score of 0.86 at a cost of 4.48 ✅ The closest-performing model: 43.48 http

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

Model ReleasesDGX agent

arXiv:2607.06118v1 Announce Type: new Abstract: As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their nav

We've usually stayed away from model comparisons but 5.6 vs Fable is a unique situation We've never had a case where the team is so complete…

Model ReleasesDGX agent

We've usually stayed away from model comparisons but 5.6 vs Fable is a unique situation We've never had a case where the team is so completely convinced on which one is better Here's the timeline of o

When Should LLMs Search? Counterfactual Supervision for Search Routing

Model ReleasesDGX agent

arXiv:2607.05752v1 Announce Type: cross Abstract: Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniformly benefici

Why does Deep Learning Improve Visual SLAM?

Model ReleasesDGX agent

arXiv:2607.06023v1 Announce Type: new Abstract: Visual SLAM is a well-established technology utilized in a wide range of real-world applications. However, its performance still degrades under challeng

Wordle 1,844 3/6 🟩⬛⬛⬛⬛ 🟨🟨🟨🟨⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This is a Wordle game result shared by Anthropic on X (formerly Twitter), showing that they solved Wordle puzzle #1,844 in 3 attempts out of 6 possible guesses. The colored squares indicate the guess

WOW! GROK 4.5 PRICES ARE NEAR OPEN SOURCE HOSTED PRICES! This will move many that have token cost fatigue in corporations. But that is only …

Model ReleasesDGX agent

WOW! GROK 4.5 PRICES ARE NEAR OPEN SOURCE HOSTED PRICES! This will move many that have token cost fatigue in corporations. But that is only half the story: it is ~4.2 more efficient using far less tim

x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability

Model ReleasesDGX agent

arXiv:2607.06114v1 Announce Type: cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (N

Xiaomi now processes more AI tokens than OpenAI. On OpenRouter, Chinese models just crossed 45% of all token volume. Anthropic is at 15.3%. …

Model ReleasesDGX agent

Xiaomi now processes more AI tokens than OpenAI. On OpenRouter, Chinese models just crossed 45% of all token volume. Anthropic is at 15.3%. OpenAI is at 7.4%. For now, the frontier is American models,

XRFormer: Multiscale Tokenization for XRF Representation Learning

Model ReleasesDGX agent

arXiv:2607.06424v1 Announce Type: new Abstract: X-ray fluorescence (XRF) spectroscopy is a key modality for material analysis in cultural heritage. However, automated learning from XRF spectra remains

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Model ReleasesDGX agent

arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children

7 Jul 2026

1 in 4 people in Toronto this week works at Cohere

Model ReleasesDGX agent

Cohere announced that approximately 25% of Toronto's working population is employed by the company during a particular week. This statement reflects Cohere's significant presence and growth as a major

20 questions for the Agentic Enterprise (and how Agent Platform can help)

Model ReleasesDGX agent

If you’re an IT leader, you might be getting a lot of questions about how to build and deploy agents. The pressure to move fast is intense, but the engineering reality is incredibly complex. Where do

A Co-Design Framework for High-Performance Jumping of a Five-Bar Monoped with Actuator Optimization

Model ReleasesDGX agent

arXiv:2604.06025v2 Announce Type: replace Abstract: The performance of legged robots depends strongly on both mechanical design and control, motivating co-design approaches that jointly optimize these

A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace

Model ReleasesDGX agent

Software-as-a-service (SaaS) is evolving into Agents-as-a-service (AaaS). Instead of isolated applications, developers are creating AI agents that interoperate using standardized open protocols such a

A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG

Model ReleasesDGX agent

arXiv:2607.03739v1 Announce Type: cross Abstract: We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning. The framework partitions rea

A Fair Benchmarking of Deep Relational Database Learning Models

Model ReleasesDGX agent

arXiv:2607.03659v1 Announce Type: cross Abstract: Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs have been eval

A Gradient Flow Perspective on Minimum MMD Estimation

Model ReleasesDGX agent

arXiv:2607.03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter

A Large-Scale Dataset and a New Method for RemoteSensing Traffic Object Segmentation

Model ReleasesDGX agent

arXiv:2607.03945v1 Announce Type: new Abstract: Remote sensing imagery plays a crucial role in evaluating regional transportation capacity. However, existing segmentation datasets often lack diversity

A Near-Linear-Time Solver for Graph p-Laplacian Semi-Supervised Learning via Continuation in p

Model ReleasesDGX agent

arXiv:2607.03503v1 Announce Type: new Abstract: Graph-based semi-supervised learning (SSL) propagates a few labels over a similarity graph by minimizing a Dirichlet-type energy. The standard quadratic

A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.04689v1 Announce Type: cross Abstract: Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded.

A Retrieval-Augmented Framework for Detecting and Resolving Pragmatic Ambiguities in Natural Language Requirements

Model ReleasesDGX agent

arXiv:2607.04436v1 Announce Type: cross Abstract: Natural language requirements (NLRs) are essential for bridging communication gaps among diverse stakeholders in software development. However, the in

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.03600v1 Announce Type: cross Abstract: Adversarial robustness in Unsupervised Domain Adaptation (UDA) remains a significant challenge due to noisy pseudo labels and inherent distributional

A Structural Interpretation of GELU and Threshold-Transmission Activations via the First-Order Loss Function

Model ReleasesDGX agent

arXiv:2607.03664v1 Announce Type: new Abstract: The Gaussian Error Linear Unit is usually motivated as the expected output of an input-dependent stochastic Bernoulli gate. This work gives a complement

A Systematic Survey on Large Language Models for Evolutionary Optimization: From Modeling to Solving

Model ReleasesDGX agent

arXiv:2509.08269v5 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly integrated with evolutionary computation to support optimization tasks. This survey primarily fo

A Technical Survey of Reinforcement Learning Techniques for Large Language Models

Model ReleasesDGX agent

arXiv:2507.04136v2 Announce Type: replace Abstract: This survey offers a comprehensive foundation on the integration of RL with language models, highlighting prominent algorithms such as Proximal Poli

Adversarial LassoNet: Robust Feature Selection via Stability-Driven Sparse Learning

Model ReleasesDGX agent

arXiv:2607.03839v1 Announce Type: new Abstract: Sparse feature selection is critical for high-dimensional machine learning, yet traditional ell_1-regularized methods are often brittle under observatio

Agent Data Injection Attacks are Realistic Threats to AI Agents

Model ReleasesDGX agent

arXiv:2607.05120v1 Announce Type: cross Abstract: AI agents act on behalf of user prompts, consuming external data and taking actions based on the agent context. Prior research on AI agent security ha

Agent-driven Long-tail Simulation for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.04331v1 Announce Type: cross Abstract: Evaluating autonomous driving systems in closed-loop settings requires realistic and interactive simulation, yet existing simulators largely rely on l

Agent Step Value: State-Transition Measurement with State-Grounded LLM Evaluators

Model ReleasesDGX agent

arXiv:2607.04419v1 Announce Type: new Abstract: Most agent evaluations collapse a multi-step trace into a final answer, a success flag, or a trajectory-level score. These aggregates obscure the diagno

AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments

Model ReleasesDGX agent

arXiv:2607.05174v1 Announce Type: new Abstract: Language agents, i.e., LLM agents, progress rapidly and are increasingly deployed in production environments. This trend underscores the urgent need for

Agentic Retrieval-Augmented Generation for Financial Document Question Answering

Model ReleasesDGX agent

arXiv:2605.05409v2 Announce Type: replace Abstract: Financial document question answering (QA) demands complex multi-step numerical reasoning over heterogeneous evidence--structured tables, textual na

AgenticPD: A Stage-Aware Agentic Framework for Physical Design QoR Optimization

Model ReleasesDGX agent

arXiv:2607.04758v1 Announce Type: new Abstract: Physical design quality-of-results~(QoR) optimization is hard and expensive. Choices made at one stage can help or hurt later stages. Each evaluation re

AgentLTL: A Trace-Verification Framework for Measuring, Enforcing, and Training Procedural Compliance in Tool-Using LLM Agents

Model ReleasesDGX agent

arXiv:2607.02599v1 Announce Type: cross Abstract: Tool-using LLM agents are usually evaluated by final-answer correctness or LLM judges. Neither captures how an answer was produced. In safety-critical

AI Wizards at EXIST 2026: Hierarchical Soft-Label Learning for Multimodal Sexism Identification in Memes

Model ReleasesDGX agent

arXiv:2607.04410v1 Announce Type: new Abstract: We present the AI Wizards submission to EXIST 2026 for multimodal sexism identification in memes. The task is composed of three, increasingly harder sub

Alibaba's Qwen models have made it an AI powerhouse, but the company has struggled to turn their global popularity into a profitable business (New York Times)

Model ReleasesDGX agent

New York Times: Alibaba's Qwen models have made it an AI powerhouse, but the company has struggled to turn their global popularity into a profitable business — The Chinese company's models have won ov

Amazon launches $25B bond sale to fund AI infrastructure

Model ReleasesDGX agent

Amazon.com Inc. is back in the bond market, this time to raise at least 25 billion for its artificial intelligence buildout, according to Bloomberg. The offering is split into eight parts, senior unse

Amortising Bayesian Experimental Design for Sequential Information Gathering in LLMs

Model ReleasesDGX agent

arXiv:2607.03426v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit strong reasoning and world-knowledge capabilities, yet often struggle to gather information effectively across th

An AI-Assisted Solution to the Signed BAR Conjecture: Uniqueness in the Harrison--Reiman Class and a Completely-S Class Obstruction

Model ReleasesDGX agent

arXiv:2607.03639v1 Announce Type: cross Abstract: For a multidimensional reflected diffusion, determining whether the associated basic adjoint relationship (BAR) uniquely characterizes the stationary

Anchored Self-Play for Code Repair

Model ReleasesDGX agent

arXiv:2607.03523v1 Announce Type: cross Abstract: Code repair is an important capability for language models (LMs): given a buggy program and unit tests, an LM must produce a fixed program that passes

AnchorSplat: Fast and Structure Consistent Detail Synthesis for Gaussian Splatting

Model ReleasesDGX agent

arXiv:2607.01290v2 Announce Type: replace Abstract: 3D Gaussian Splatting (3DGS) has emerged as a powerful representation for high-fidelity rendering. However, existing assets often suffer from qualit

AnchorVLA: Bridging Discrete Decisions and Continuous Trajectories for Vision-Language-Action Planning

Model ReleasesDGX agent

arXiv:2607.03182v1 Announce Type: cross Abstract: Autonomous driving planning requires translating navigation intent, traffic rules, dynamic interactions, and language instructions into executable con

Another day another big public tech co saying they’re using open source AI models extensively

Model ReleasesDGX agent

Another day another big public tech co saying they’re using open source AI models extensively With our internal coding benchmark, we're able to confidently introduce open-weight models into our AI cod

Anthropic brings Cowork out of the desktop and onto web and mobile

Model ReleasesDGX agent

Anthropic PBC today announced it’s bringing its Claude Cowork agentic artificial intelligence assistant to mobile and the web, breaking it free from the desktop. Cowork allows users to harness the com

Anthropic expands Claude Cowork to web and mobile in beta for Max plan subscribers, and says 90%+ of Cowork usage is unrelated to software development (David Gewirtz/ZDNET)

Model ReleasesDGX agent

David Gewirtz / ZDNET: Anthropic expands Claude Cowork to web and mobile in beta for Max plan subscribers, and says 90%+ of Cowork usage is unrelated to software development — ZDNET's key takeaways —

Anthropic is launching Claude Cowork on mobile and web

Model ReleasesDGX agent

Starting Tuesday, Anthropic's Claude Cowork AI platform will be available on mobile and web for the first time. The expanded access is rolling out first to Max subscribers and coming to Claude users o

Anthropic releasing open source demos built on top of Qwen (with @neuronpedia) is not something that I was expecting

Model ReleasesDGX agent

Anthropic releasing open source demos built on top of Qwen (with @neuronpedia) is not something that I was expecting We also partnered with Neuronpedia to create an interactive demo of our methods on

APeB: Benchmarking Personalization Ability of Large Language Model Agents

Model ReleasesDGX agent

arXiv:2607.03162v1 Announce Type: new Abstract: LLM-powered agents struggle with personalization when users issue raw, underspecified queries. In this setting, agents must infer latent intent, extract

← Previous
1…96979899100…377
Next →