AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,570
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,566
  • Research19,194
  • Safety12,816
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,570Total entries
1Added by human
84,569Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,566 results
Model Releases

Spider 2.0-AIFunc: Extending Real-World Text-to-SQL to AI-Native SQL Workflows

DGX agent

arXiv:2607.06229v1 Announce Type: cross Abstract: Major cloud data platforms now expose large language model capabilities as native SQL functions, enabling analysts to perform classification, filterin

model-releasesarxiv-cs-ai
8 Jul 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Statistically Meaningful Geometry and Gauge Symmetry Breaking: A Geometric Foundation for Scientific Discovery and Intelligence Emergence

DGX agent

arXiv:2607.05436v1 Announce Type: new Abstract: The rapid scaling of over-parameterized machine learning architectures, particularly LLMs, raises a profound crisis: do these systems exhibit genuine in

model-releasesarxiv-cs-lg
8 Jul 2026
Model Releases

StepShield: When, Not Whether to Intervene on Rogue Agents

DGX agent

arXiv:2601.22136v2 Announce Type: replace-cross Abstract: Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We in

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

DGX agent

arXiv:2607.06012v1 Announce Type: new Abstract: Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.e., legal status, structural

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

super interesting - and a reminder that language models model language and not ideas per se. also a stark example of how easily LLMs can (as…

DGX agent

super interesting - and a reminder that language models model language and not ideas per se. also a stark example of how easily LLMs can (as a consequence) be influenced by propaganda. Asking ChatGPT

model-releasesgary-marcus--x
8 Jul 2026
Model Releases

Superhuman launches Docs, merging writing, AI and data for document collaboration

DGX agent

Superhuman Inc., the company formerly known as Grammarly, today announced the launch of Docs, a product that enables multiple users to collaborate on writing using artificial intelligence. Docs contai

model-releasessiliconangle
8 Jul 2026
Model Releases

Supervised Reward Inference

DGX agent

arXiv:2502.18447v2 Announce Type: replace Abstract: Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models. However, humans o

model-releasesarxiv-cs-lg
8 Jul 2026
Model Releases

The anthropomorphization of this one is off the charts. Why does @AnthropicAI do this? It is intellectually lazy. Unscientific.

DGX agent

The anthropomorphization of this one is off the charts. Why does @AnthropicAI do this? It is intellectually lazy. Unscientific. New Anthropic research: A global workspace in language models. Of everyt

model-releasesgary-marcus--x
8 Jul 2026
Model Releases

The Balkanization of Execution-Security Research for AI Coding Agents: Isolation, Access Control, and Time-of-Check-to-Time-of-Use Vulnerabilities

DGX agent

arXiv:2607.05743v1 Announce Type: cross Abstract: AI coding agents now read repositories, call tools, and execute shell commands with limited human oversight, and a fast-growing body of work studies w

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

The Granularity Paradox: How Temporal Disaggregation Inflates In-Sample Fit and Compounds Out-of-Sample Error

DGX agent

arXiv:2607.05450v1 Announce Type: cross Abstract: This paper explores the 'Granularity Paradox' in time-series forecasting, wherein finer temporal disaggregation (e.g., Monthly to Weekly/Daily) improv

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

The next generation of ChatGPT Voice is here. Livestream starts at 10am PT. https://openai.com/live/

DGX agent

OpenAI announced a livestream event showcasing the next generation of ChatGPT Voice, scheduled to begin at 10am PT. The announcement was made via OpenAI's official X (formerly Twitter) account, direct

model-releasesopenai--x
8 Jul 2026
Model Releases

The relationship between reasoning and performance in large language models--o3 (mini) thinks harder, not longer

DGX agent

arXiv:2502.15631v2 Announce Type: replace-cross Abstract: Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning.

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

The yes-no bias of large language models reflects answer order and wording, not shifts in moral judgment

DGX agent

arXiv:2607.05552v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly issue judgments read as binary verdicts, and a growing literature reports such judgments shifting under logi

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

These new models are rolling out to everyone over the next few days, starting today. They’ll be available in ChatGPT across iOS, Android, an…

DGX agent

These new models are rolling out to everyone over the next few days, starting today. They’ll be available in ChatGPT across iOS, Android, and web. Coming soon to the API. Just tap the Voice button to

model-releasesopenai--x
8 Jul 2026
Model Releases

Think Before You Grid-Search: Floor-First Triage for LLM Serving

DGX agent

arXiv:2607.05876v1 Announce Type: cross Abstract: LLM serving optimization typically benchmarks many configurations and reaches for heavy profilers when latency targets are missed. We argue for the re

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

ThorArena: Benchmarking Humanoid Physical Interaction with Human Motion-Force Demonstrations

DGX agent

arXiv:2607.06052v1 Announce Type: new Abstract: Humanoid robots are increasingly expected to perform contact-rich tasks that require not only accurate whole-body motion but also robust physical intera

model-releasesarxiv-cs-ro
8 Jul 2026
Model Releases

To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software eng…

DGX agent

To audit SWE-Bench Pro, we used model-based investigator agents alongside independent reviews from five independent experienced software engineers. That helped us examine tasks at scale while keeping

model-releasesopenai--x
8 Jul 2026
Model Releases

Tomorrow at 10am PT I'm hosting a live walkthrough of how we progressed from single-player Claude Code to multi-player Claude Tag. Then, we'…

DGX agent

Tomorrow at 10am PT I'm hosting a live walkthrough of how we progressed from single-player Claude Code to multi-player Claude Tag. Then, we're going deep on how Claude Tag actually works. AI used to f

model-releasesboris-cherny--x
8 Jul 2026
Model Releases

Trained entirely in simulation: ~400,000 trajectories across 6,000 scenes. A prefix-caching recipe cuts training tokens by 22×, turning mont…

DGX agent

Trained entirely in simulation: ~400,000 trajectories across 6,000 scenes. A prefix-caching recipe cuts training tokens by 22×, turning months-long runs into days. Online RL (CISPO) pushes success rat

model-releasesmistral-ai--x
8 Jul 2026
Model Releases

Transformers converge to invariant algorithmic cores

DGX agent

arXiv:2602.22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function. Studying any single trained neural n

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

UI2App: Benchmarking Visual Interaction Inference in Executable Web Application Generation

DGX agent

arXiv:2607.06306v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex pro

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

UniLM-Nav: A Unified Framework for Zero-Shot Last-Mile Navigation

DGX agent

arXiv:2607.06537v1 Announce Type: new Abstract: Mobile manipulation requires a robot to navigate to a target object or receptacle and then perform intended manipulation. However, reaching the vicinity

model-releasesarxiv-cs-ro
8 Jul 2026
Model Releases

VendorBench-100: A Unified Cross-Paradigm Benchmark for Deepfake Image Detection

DGX agent

arXiv:2607.06254v1 Announce Type: cross Abstract: Deepfake image detection is currently served by three fundamentally different paradigms: commercial APIs, zero-shot vision-language models (LLMs), and

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Verification of Dynamic Holographic Behavior in Identity Documents

DGX agent

arXiv:2607.06466v1 Announce Type: new Abstract: This paper addresses the remote verification of the authenticity of Optically Variable Devices (commonly known as holograms) on identity documents. Typi

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models

DGX agent

arXiv:2510.13808v2 Announce Type: replace Abstract: Large Vision Language Models (VLMs) excel at general visual reasoning but experience significant performance degradation when deployed in novel doma

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capabil…

DGX agent

We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and ar

model-releasesopenai--x
8 Jul 2026
Model Releases

We need AI model selection to be MUCH easier ASAP. I want proactive flags from my AI systems suggesting models. I want my AI harness to say …

DGX agent

We need AI model selection to be MUCH easier ASAP. I want proactive flags from my AI systems suggesting models. I want my AI harness to say 'hey allie, my girl, you keep asking for bar recommendations

model-releasesallie-k--miller--x
8 Jul 2026
Model Releases

We tuned the harness for @NVIDIAAI Nemotron 3 Ultra. Benchmark-leading performance. 10x lower inference costs. ✅ An aggregate score of 0.86 …

DGX agent

We tuned the harness for @NVIDIAAI Nemotron 3 Ultra. Benchmark-leading performance. 10x lower inference costs. ✅ An aggregate score of 0.86 at a cost of 4.48 ✅ The closest-performing model: 43.48 http

model-releasesharrison-chase--x
8 Jul 2026
Model Releases

WebRetriever: A Large-Scale Comprehensive Benchmark for Efficient Web Agent Evaluation

DGX agent

arXiv:2607.06118v1 Announce Type: new Abstract: As web agents increasingly demonstrate capabilities in automated task execution, the development of robust evaluation frameworks for assessing their nav

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

We've usually stayed away from model comparisons but 5.6 vs Fable is a unique situation We've never had a case where the team is so complete…

DGX agent

We've usually stayed away from model comparisons but 5.6 vs Fable is a unique situation We've never had a case where the team is so completely convinced on which one is better Here's the timeline of o

model-releasessam-altman--x
8 Jul 2026
Model Releases

When Should LLMs Search? Counterfactual Supervision for Search Routing

DGX agent

arXiv:2607.05752v1 Announce Type: cross Abstract: Search-augmented language models can use external evidence to compensate for limitations in parametric knowledge, but search is not uniformly benefici

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Why does Deep Learning Improve Visual SLAM?

DGX agent

arXiv:2607.06023v1 Announce Type: new Abstract: Visual SLAM is a well-established technology utilized in a wide range of real-world applications. However, its performance still degrades under challeng

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Wordle 1,844 3/6 🟩⬛⬛⬛⬛ 🟨🟨🟨🟨⬛ 🟩🟩🟩🟩🟩

DGX agent

This is a Wordle game result shared by Anthropic on X (formerly Twitter), showing that they solved Wordle puzzle #1,844 in 3 attempts out of 6 possible guesses. The colored squares indicate the guess

model-releasesanthropic--x
8 Jul 2026
Model Releases

WOW! GROK 4.5 PRICES ARE NEAR OPEN SOURCE HOSTED PRICES! This will move many that have token cost fatigue in corporations. But that is only …

DGX agent

WOW! GROK 4.5 PRICES ARE NEAR OPEN SOURCE HOSTED PRICES! This will move many that have token cost fatigue in corporations. But that is only half the story: it is ~4.2 more efficient using far less tim

model-releaseselon-musk--x
8 Jul 2026
Model Releases

x-Prediction Is All You Need:Training-Free Accelerated Generation via Endpoint Decodability

DGX agent

arXiv:2607.06114v1 Announce Type: cross Abstract: Diffusion and flow matching models generate high-quality samples, but their ODE samplers often need tens to hundreds of neural function evaluations (N

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Xiaomi now processes more AI tokens than OpenAI. On OpenRouter, Chinese models just crossed 45% of all token volume. Anthropic is at 15.3%. …

DGX agent

Xiaomi now processes more AI tokens than OpenAI. On OpenRouter, Chinese models just crossed 45% of all token volume. Anthropic is at 15.3%. OpenAI is at 7.4%. For now, the frontier is American models,

model-releasesallie-k--miller--x
8 Jul 2026
Model Releases

XRFormer: Multiscale Tokenization for XRF Representation Learning

DGX agent

arXiv:2607.06424v1 Announce Type: new Abstract: X-ray fluorescence (XRF) spectroscopy is a key modality for material analysis in cultural heritage. However, automated learning from XRF spectra remains

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

DGX agent

arXiv:2502.09696v3 Announce Type: replace Abstract: Large Multimodal Models (LMMs) exhibit shortfalls when interpreting images and, by some measures, have poorer spatial cognition than young children

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

1 in 4 people in Toronto this week works at Cohere

DGX agent

Cohere announced that approximately 25% of Toronto's working population is employed by the company during a particular week. This statement reflects Cohere's significant presence and growth as a major

model-releasescohere--x
7 Jul 2026
Model Releases

20 questions for the Agentic Enterprise (and how Agent Platform can help)

DGX agent

If you’re an IT leader, you might be getting a lot of questions about how to build and deploy agents. The pressure to move fast is intense, but the engineering reality is incredibly complex. Where do

model-releasesgoogle-cloud-ai
7 Jul 2026
Model Releases

A Co-Design Framework for High-Performance Jumping of a Five-Bar Monoped with Actuator Optimization

DGX agent

arXiv:2604.06025v2 Announce Type: replace Abstract: The performance of legged robots depends strongly on both mechanical design and control, motivating co-design approaches that jointly optimize these

model-releasesarxiv-cs-ro
7 Jul 2026
Model Releases

A developer's guide to publishing agents in Gemini Enterprise and Google Cloud Marketplace

DGX agent

Software-as-a-service (SaaS) is evolving into Agents-as-a-service (AaaS). Instead of isolated applications, developers are creating AI agents that interoperate using standardized open protocols such a

model-releasesgoogle-cloud-ai
7 Jul 2026
Model Releases

A Failure-Mode Benchmark for Polymorphic Sybil Poisoning in RAG

DGX agent

arXiv:2607.03739v1 Announce Type: cross Abstract: We release a benchmark and failure-mode-aware evaluation framework for grounded QA under coordinated retrieval poisoning. The framework partitions rea

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

A Fair Benchmarking of Deep Relational Database Learning Models

DGX agent

arXiv:2607.03659v1 Announce Type: cross Abstract: Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs have been eval

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

A Gradient Flow Perspective on Minimum MMD Estimation

DGX agent

arXiv:2607.03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

A Large-Scale Dataset and a New Method for RemoteSensing Traffic Object Segmentation

DGX agent

arXiv:2607.03945v1 Announce Type: new Abstract: Remote sensing imagery plays a crucial role in evaluating regional transportation capacity. However, existing segmentation datasets often lack diversity

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

A Near-Linear-Time Solver for Graph p-Laplacian Semi-Supervised Learning via Continuation in p

DGX agent

arXiv:2607.03503v1 Announce Type: new Abstract: Graph-based semi-supervised learning (SSL) propagates a few labels over a similarity graph by minimizing a Dirichlet-type energy. The standard quadratic

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

A Reliable Context-Aware and Temporal Planning Framework for Autonomous Driving

DGX agent

arXiv:2607.04689v1 Announce Type: cross Abstract: Safe operation of autonomous vehicles in dense urban traffic depends on perception and planning that remain reliable when onboard sensing is degraded.

model-releasesarxiv-cs-cv
7 Jul 2026
← Previous
1…119120121122123…471
Next →