AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,561 results
7 Jul 2026

Triple-Phase Multimodal Knowledge Aggregation Framework for Microbial Keratitis Subtype Diagnosis on Slit-Lamp Photography

Model ReleasesDGX agent

arXiv:2607.03740v1 Announce Type: cross Abstract: Microbial keratitis requires rapid pathogen identification to guide treatment, but culture- and PCR-based diagnostics are slow and resource-intensive.

TSP with Predictions: Heatmap to Tour with Provable Guarantees

Model ReleasesDGX agent

arXiv:2607.03791v1 Announce Type: cross Abstract: The Traveling Salesperson Problem (TSP) has long served as a benchmark for evaluating the strength of optimization techniques in the classical theory

U-Joint CAAMS: Experimental Evaluation of a Universal-Joint Continuum Manipulator for Aerial Manipulation


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2607.03321v1 Announce Type: new Abstract: Continuum manipulators mounted on multi-rotor UAVs enable compliant aerial manipulation, but payloads and propeller downwash amplify out-of-plane bendin

Unbiased Alignment for Large Language Models with Noisy Preferences

Model ReleasesDGX agent

arXiv:2607.03248v1 Announce Type: cross Abstract: The alignment of large language models with human preferences is commonly achieved through Reinforcement Learning from Human Feedback or Direct Prefer

Uncertainty-aware damage identification in short-span bridges via physics-informed variational autoencoder

Model ReleasesDGX agent

arXiv:2607.05025v1 Announce Type: new Abstract: Vibration-based damage identification in civil infrastructure is a challenging, ill-posed inverse problem due to measurement noise, sparse sensor arrays

Unified Audio Intelligence Without Regressing on Text Intelligence

Model ReleasesDGX agent

arXiv:2607.05196v1 Announce Type: cross Abstract: Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

Model ReleasesDGX agent

arXiv:2603.24690v2 Announce Type: replace Abstract: In-context learning (ICL) enables fast task adaptation from demonstrations without per-task parameter updates but remains highly sensitive to exampl

UniVideo: Unified Understanding, Generation, and Editing for Videos

Model ReleasesDGX agent

arXiv:2510.08377v4 Announce Type: replace Abstract: Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain.

Unsupervised Behavioral Compression: Learning Low-Dimensional Policy Manifolds through State-Occupancy Matching

Model ReleasesDGX agent

arXiv:2603.27044v3 Announce Type: replace-cross Abstract: Deep Reinforcement Learning (DRL) is widely recognized as sample-inefficient, a limitation attributable in part to the high dimensionality and

URSA: Chemistry-Aware Benchmark for Utilitarian Retrosynthesis Assessment

Model ReleasesDGX agent

arXiv:2607.04688v1 Announce Type: cross Abstract: Synthesis planning aiming to find pathways of reactions for a target molecule is one of the most important and challenging tasks in drug discovery. Re

Variable Bit-width Quantization: Learning Per-Group Precision for 'Bigger-but-Smaller' Language Models

Model ReleasesDGX agent

arXiv:2607.02893v1 Announce Type: cross Abstract: Low-bit quantization shrinks language models but treats precision as a single global hyper-parameter: every weight uses the same bit-width. We introdu

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

Model ReleasesDGX agent

arXiv:2510.11098v5 Announce Type: replace-cross Abstract: Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks r

VERITAS: Towards a General-Purpose Replication Tool for Scientific Research

Model ReleasesDGX agent

arXiv:2607.02931v1 Announce Type: new Abstract: AI tools are accelerating scientific publication while the systems that review it struggle to keep up, and independent verification of published researc

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2607.02927v1 Announce Type: cross Abstract: Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (V

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

Model ReleasesDGX agent

arXiv:2511.20272v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of

VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Model ReleasesDGX agent

arXiv:2407.11691v5 Announce Type: replace Abstract: We present VLMEvalKit: an open-source toolkit for evaluating large multi-modality models based on PyTorch. The toolkit aims to provide a user-friend

Walrus: A Cross-Domain Foundation Model for Continuum Dynamics

Model ReleasesDGX agent

arXiv:2511.15684v2 Announce Type: replace-cross Abstract: Foundation models have transformed machine learning for language and vision, but achieving comparable impact in physical simulation remains a

We're extending access to Claude Fable 5 on all paid plans through July 12.

Model ReleasesDGX agent

Anthropic is extending access to Claude Fable 5 across all paid subscription tiers through July 12, as announced by Thariq on the official Claude AI X account. This indicates a temporary or promotiona

What’s New in Microsoft Foundry | June 2026

Model ReleasesDGX agent

Claude is now generally available in Microsoft Foundry. Here's everything else that shipped between Build 2026 and the end of June — autopilot agents, expanded Toolboxes and Routines, Agent Optimizer'

When Aggregate Alignment Misleads: Auditing Policy Repair Without Per-State Expert Actions

Model ReleasesDGX agent

arXiv:2607.03386v1 Announce Type: new Abstract: Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert ac

When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents

Model ReleasesDGX agent

arXiv:2607.05189v1 Announce Type: cross Abstract: Persistent personal agents combine long-term memory with access to users' external environments, enabling personalized foreground assistance and proac

When Do Foundation Models Pay Off? A Break-Even Analysis of Pretrained Time Series Forecasters

Model ReleasesDGX agent

arXiv:2607.04919v1 Announce Type: new Abstract: Deploying a time series foundation model requires GPU infrastructure, engineering overhead, and carries no guarantee of improvement over XGBoost. We pro

When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions

Model ReleasesDGX agent

arXiv:2607.04731v1 Announce Type: new Abstract: Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising pr

When is a System Discoverable from Data? Discovery Requires Chaos

Model ReleasesDGX agent

arXiv:2511.08860v2 Announce Type: replace-cross Abstract: The deep learning revolution has spurred a rise in advances of using AI in sciences. Within physical sciences the main focus has been on disco

When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On

Model ReleasesDGX agent

arXiv:2603.05659v3 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and Rubrics as Rewards (RaR) have driven strong gains in domains with clear correctness

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

Model ReleasesDGX agent

arXiv:2607.03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natu

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

Model ReleasesDGX agent

arXiv:2510.19186v3 Announce Type: replace Abstract: Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and to

Which Algorithm Specification Formats Help Language Models Implement Machine Learning Algorithms?

Model ReleasesDGX agent

arXiv:2607.03158v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to implement algorithms from research manuscripts, but papers often leave implementation choices im

Why Open Data Matters | Nemotron Labs https://x.com/i/broadcasts/1DxLddoaAOoxm

Model ReleasesDGX agent

This broadcast likely discusses the importance and benefits of open data in AI development and machine learning, covering topics such as accessibility, transparency, democratization of AI technology,

Wordle 1,843 6/6 ⬛🟨⬛⬛⬛ ⬛⬛⬛⬛🟨 🟩🟩⬛⬛⬛ 🟩🟩⬛⬛🟩 🟩🟩⬛⬛🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This is a Wordle game result shared by Anthropic on X (Twitter), showing a player who solved puzzle #1,843 on their sixth and final attempt. The colored emoji squares document the progression of guess

WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling

Model ReleasesDGX agent

arXiv:2607.03461v1 Announce Type: new Abstract: World models aim to capture environment dynamics in ways that support perception, reasoning, and action, and have recently become a central direction in

Wrong Before Right: Late Rescue and Interface Failure in Aligned Language Models

Model ReleasesDGX agent

arXiv:2607.04640v1 Announce Type: new Abstract: We study how correctness is assembled inside aligned language models, not only whether the final answer is right. Using layer-wise difference-in-differe

WSA_1: a 3D-Centric World-Spatial-Action Model for Generalizable Robot Control

Model ReleasesDGX agent

arXiv:2607.03941v1 Announce Type: new Abstract: Recent advances in embodied AI have established robot foundation models (RFMs) as the dominant approach for generalist robotic systems to date. By lever

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

Model ReleasesDGX agent

arXiv:2607.03562v1 Announce Type: new Abstract: As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limi

XS-VLA: Coupling Coarse-grained Spatial Distillation with Latent Flow Matching for Lightweight Robotic Control

Model ReleasesDGX agent

arXiv:2607.04171v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong multimodal understanding and spatial grounding, but their computational cost limits real-time r

6 Jul 2026

128 GB of memory is nice, but you can get started with local agentic AI workflows with much less. By connecting gemma 4 in @lmstudio to MATL…

Model ReleasesDGX agent

128 GB of memory is nice, but you can get started with local agentic AI workflows with much less. By connecting gemma 4 in @lmstudio to MATLAB MCP Server, you can run a local AI model that uses MATLAB

By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.

Model ReleasesDGX agent

This post discusses how observing Claude's internal activation patterns (J-space) reveals the model's hidden reasoning processes, such as detecting code bugs and analyzing images, even when these step

Claude Fable running ComfyUI workflows through the Comfy MCP is a cheat code. → Claude pulled the shot from the web, auto-detected the scene…

Model ReleasesDGX agent

Claude Fable running ComfyUI workflows through the Comfy MCP is a cheat code. → Claude pulled the shot from the web, auto-detected the scene cuts + trimmed it → ran my saved depthanything v3 + openpos

Claude Opus 4.8 and Sonnet 5 seem worse at tool calls than older models, likely due to post-training that assumes Claude Code-like harnesses as targets (Armin Ronacher/Armin Ronacher's Thoughts and Writings)

Model ReleasesDGX agent

Armin Ronacher / Armin Ronacher's Thoughts and Writings: Claude Opus 4.8 and Sonnet 5 seem worse at tool calls than older models, likely due to post-training that assumes Claude Code-like harnesses as

During a Bloomberg interview, Yann LeCun (@ylecun ) explains why LLMs are limited in terms of real-world intelligence during a Bloomberg int…

Model ReleasesDGX agent

During a Bloomberg interview, Yann LeCun (@ylecun ) explains why LLMs are limited in terms of real-world intelligence during a Bloomberg interview. 'Language is a very approximate, reduced, quantized,

'Ghost memory' is a real problem with agents. You might have seen the issue where a long-running agent still confidently repeats a user fact…

Model ReleasesDGX agent

'Ghost memory' is a real problem with agents. You might have seen the issue where a long-running agent still confidently repeats a user fact that stopped being true weeks ago? New research names the f

Interesting stuff. And the visualization at the end is worth trying: https://www.neuronpedia.org/qwen3.6-27b/jlens

Model ReleasesDGX agent

Interesting stuff. And the visualization at the end is worth trying: https://www.neuronpedia.org/qwen3.6-27b/jlens New Anthropic research: A global workspace in language models. Of everything happenin

Less than 2 years later, with Fable: 'simulate an encounter between a mind flayer and a drow warrior of equal CR. Set up initial stats then …

Model ReleasesDGX agent

Less than 2 years later, with Fable: 'simulate an encounter between a mind flayer and a drow warrior of equal CR. Set up initial stats then simulate each move, including initiative, rolling dice as ne

Must-read research by Anthropic. Here is the simple explanation and why this is a big deal. We suspect LLMs perform 'internal reasoning'. Bu…

Model ReleasesDGX agent

Must-read research by Anthropic. Here is the simple explanation and why this is a big deal. We suspect LLMs perform 'internal reasoning'. But little is known or do good methods exist to understand it.

Not to excuse myself but for both OpenAI and Claude, what settings are global, which ones apply to local machines, and which apply to web an…

Model ReleasesDGX agent

Ethan Mollick discusses the configuration scope differences between OpenAI and Claude AI systems, distinguishing between global settings that apply across all uses, local machine-specific settings, an

One of the only times I remind people I have a PhD in computational neuroscience is when people without a neuroscience background say their …

Model ReleasesDGX agent

One of the only times I remind people I have a PhD in computational neuroscience is when people without a neuroscience background say their model works 'like the brain.' In these cases, I put on my ne

Our 17th Transporter rideshare mission is targeted to launch tomorrow from California and will deliver 81 payloads to orbit → http://spacex.…

Model ReleasesDGX agent

SpaceX's 17th Transporter rideshare mission is scheduled to launch from California, carrying 81 payloads to orbit. Transporter missions are SpaceX's dedicated rideshare service that launches multiple

Reporting benchmark results as a scalar number, e.g. '75% on XYZ' is completely meaningless at this point. You should always report efficien…

Model ReleasesDGX agent

Reporting benchmark results as a single scalar percentage is insufficient for meaningful evaluation of model performance. Comprehensive benchmark reporting should include efficiency metrics alongside

Shift into high gear with agents: Securing the software-defined vehicle

Model ReleasesDGX agent

The automotive industry is at a pivotal crossroads as it hits the gas on adopting new technology. The era of the traditional connected vehicle has shifted into the age of the software-defined vehicle

SK Hynix launches a US share sale on the Nasdaq to raise ~$28B, selling 17.79M new shares; the final listing price is due Thursday before trading starts Friday (Reuters)

Model ReleasesDGX agent

Reuters: SK Hynix launches a US share sale on the Nasdaq to raise ~28B, selling 17.79M new shares; the final listing price is due Thursday before trading starts Friday — South Korean chipmaker SK Hyni

so much for recursive self improvement, to the degree that it requires scientific taste

Model ReleasesDGX agent

so much for recursive self improvement, to the degree that it requires scientific taste the other thing im noticing while working on my research projects is how limited these models are GPT-5.5-xhigh

sqlite-utils 4.0rc3

Model ReleasesDGX agent

Release: sqlite-utils 4.0rc3 I hoped to release sqlite-utils 4.0 stable this weekend, but as I worked through the backlog of issues and PRs with a combination of Claude Fable 5 and GPT-5.5 the changel

Streaming benchmark and recommendation results to MLflow with Amazon SageMaker AI

Model ReleasesDGX agent

In this post, you learn how to use the new MLflow integration with Amazon SageMaker AI optimized inference recommendation jobs and Amazon SageMaker AI benchmark jobs to automatically stream experiment

tencent/Hy3

Model ReleasesDGX agent

tencent/Hy3 New Apache 2.0 licensed model from Tencent in China: Hy3 is a 295B-parameter Mixture-of-Experts (MoE) model with 21B active parameters and 3.8B MTP layer parameters, developed by the Tence

The Nemotron family just passed 100M downloads! Huge thank you to the community building with us and showing what’s possible with open model…

Model ReleasesDGX agent

The Nemotron model family from NVIDIA has reached 100 million downloads, marking a significant milestone for the open-source AI model community. The achievement reflects growing adoption and collabora

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety resea…

Model ReleasesDGX agent

This is our first time telling the story of how we first built and launched Claude Code, starting with its origins in Anthropic safety research. So much more to do. We are 1% done. We've put together

“This study cuts through the optimism surrounding medical AI by showing how easily benchmark success can be mistaken for real readiness. In …

Model ReleasesDGX agent

“This study cuts through the optimism surrounding medical AI by showing how easily benchmark success can be mistaken for real readiness. In medical AI, impressive scores are clearly not the same as tr

We released our in-house Japanese, English, and Chinese translation tool, Sakana Translate! Try it→ https://translate.sakana.ai 🐟 • Transla…

Model ReleasesDGX agent

We released our in-house Japanese, English, and Chinese translation tool, Sakana Translate! Try it→ https://translate.sakana.ai 🐟 • Translate: Handles long text in real time • Proofread: Tone and phra

When it comes to Chess, @viditchess is an absolute legend. Recently, he entered the Hermes Agent Accelerated Business Hackathon to see who c…

Model ReleasesDGX agent

When it comes to Chess, @viditchess is an absolute legend. Recently, he entered the Hermes Agent Accelerated Business Hackathon to see who could build the best agent using Nemotron. Here's what he sub

With our internal coding benchmark, we're able to confidently introduce open-weight models into our AI code reviewer w/o degrading code qual…

Model ReleasesDGX agent

With our internal coding benchmark, we're able to confidently introduce open-weight models into our AI code reviewer w/o degrading code quality. Have the frontier model (Fable) to the hardest work, de

← Previous
1…103104105106107…377
Next →