AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,553 results
6 May 2026

Valley3: Scaling Omni Foundation Models for E-commerce

Model ReleasesDGX agent

arXiv:2605.01278v1 Announce Type: new Abstract: In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understandi

Vanishing L2 regularization for the softmax Multi Armed Bandit

Model ReleasesDGX agent

arXiv:2605.03752v1 Announce Type: new Abstract: Multi Armed Bandit (MAB) algorithms are a cornerstone of reinforcement learning and have been studied both theoretically and numerically. One of the mos

@vasuman We've been saying this for months. The best compliment you can give an AI system is that it behaves exactly as expected. Every time…

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

@vasuman We've been saying this for months. The best compliment you can give an AI system is that it behaves exactly as expected. Every time. We built a campaign around it. It's called Boring AI: http

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

Model ReleasesDGX agent

arXiv:2605.03276v1 Announce Type: new Abstract: Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage i

very fun to collab with @harvey on their Long Horizon Legal Agent Benchmark. We need more industry specific benchmarks, and Harvey is paving…

Model ReleasesDGX agent

Harrison Chase expresses enthusiasm about collaborating with Harvey on their Long Horizon Legal Agent Benchmark, highlighting the value of developing industry-specific benchmarks for AI evaluation. Th

💫Very happy to release NeuralBench, to benchmark Neuro AI models and datasets in the open! 🧵Thread, 💻Code, 📝White Paper below:

Model ReleasesDGX agent

💫Very happy to release NeuralBench, to benchmark Neuro AI models and datasets in the open! 🧵Thread, 💻Code, 📝White Paper below: 🧠 Introducing NeuralBench: a unified, open-source framework to benchmark

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

Model ReleasesDGX agent

arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete '

Vibe coding and agentic engineering are getting closer than I'd like

Model ReleasesDGX agent

I recently talked with Joseph Ruscio about AI coding tools for Heavybit's High Leverage podcast: Ep. #9, The AI Coding Paradigm Shift with Simon Willison. Here are some of my highlights, including my

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.01449v1 Announce Type: cross Abstract: Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, sug

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.03351v1 Announce Type: new Abstract: Video vision-language models (VLMs) keep paying for visual state the stream already told us was stable. The factory wall did not move, but most VLM pipe

We also announced a compute partnership with SpaceX: 300+ MW of new capacity and 220K NVIDIA GPUs online within the month, all powering Clau…

Model ReleasesDGX agent

Anthropic announced a major compute partnership with SpaceX involving over 300 MW of new computational capacity and 220,000 NVIDIA GPUs to be brought online within one month, dedicated to powering Cla

we're continuing to see clear examples where a model's harness is a major determinant of overall performance. with the same model, running o…

Model ReleasesDGX agent

we're continuing to see clear examples where a model's harness is a major determinant of overall performance. with the same model, running on same task, it's easy to observe very different scores depe

We're winding back our peak hours limit reduction and doubling 5 hour limits. Excited to partner with SpaceX to bring you more compute and w…

Model ReleasesDGX agent

We're winding back our peak hours limit reduction and doubling 5 hour limits. Excited to partner with SpaceX to bring you more compute and we'll keep pushing to bring you the best coding agent in the

We’ve agreed to a partnership with @SpaceX that will substantially increase our compute capacity. This, along with our other recent compute …

Model ReleasesDGX agent

We’ve agreed to a partnership with @SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limi

We’ve been working closely with the @harvey team on the launch of the Legal Agent Benchmark, a product focused on evaluating how open-weight…

Model ReleasesDGX agent

We’ve been working closely with the @harvey team on the launch of the Legal Agent Benchmark, a product focused on evaluating how open-weight models perform on long-horizon, real-world legal tasks. Che

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-paramet…

Model ReleasesDGX agent

We’ve developed our own inference engine Runtime-Optimized Serving Engine (ROSE) to serve models ranging from embeddings to trillion-parameter LLMs. With CuTeDSL integrated into our inference engine,

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powe…

Model ReleasesDGX agent

What if you could extract text from any photo on your phone? We built LlamaParse Mobile, an @expo + @reactnative app for iOS & Android, powered by the LlamaParse TypeScript SDK 📱 Three steps, that’s i

What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models

Model ReleasesDGX agent

arXiv:2603.12799v2 Announce Type: replace Abstract: Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and chal

What's new in IAM: Security, governance, and runtime defense

Model ReleasesDGX agent

The AI era demands a fundamental shift in security, and that includes identity and access management (IAM). Traditional controls simply aren’t built for autonomous AI agents that interact with sensiti

When Alignment Isn't Enough: Response-Path Attacks on LLM Agents

Model ReleasesDGX agent

arXiv:2605.02187v1 Announce Type: cross Abstract: Bring-Your-Own-Key (BYOK) agent architectures let users route LLM traffic through third-party relays, creating a critical integrity gap: a malicious r

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

Model ReleasesDGX agent

arXiv:2605.03096v1 Announce Type: cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

Model ReleasesDGX agent

arXiv:2605.02915v1 Announce Type: new Abstract: Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its pr

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2605.02463v1 Announce Type: cross Abstract: Multi-agent LLM systems are increasingly used to solve complex tasks through decomposition, debate, specialization, and ensemble reasoning. However, t

Where to Bind Matters: Hebbian Fast Weights in Vision Transformers for Few-Shot Character Recognition

Model ReleasesDGX agent

arXiv:2605.02920v1 Announce Type: cross Abstract: Standard transformer architectures learn fixed slow-weight representations during training and lack mechanisms for rapid adaptation within an episode.

Wordle 1,781 3/6 ⬛🟨🟨⬛⬛ 🟩⬛⬛⬛⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

Anthropic shared their Wordle game result for puzzle #1,781, solved in 3 attempts with a final correct answer shown by five green squares (all letters in correct positions). The emoji grid represents

Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

Model ReleasesDGX agent

arXiv:2605.03596v1 Announce Type: cross Abstract: Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

Model ReleasesDGX agent

arXiv:2605.03475v1 Announce Type: new Abstract: Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal t

xAI and SpaceXAI have just made Colossus 1 available to Anthropic to support Claude. This means more than 220,000 NVIDIA GPUs in one of the …

Model ReleasesDGX agent

xAI and SpaceXAI have just made Colossus 1 available to Anthropic to support Claude. This means more than 220,000 NVIDIA GPUs in one of the world’s largest and fastest-built AI superclusters are now h

5 May 2026

A Closed-Form Persistence-Landmark Pipeline for Certified Point-Cloud and Graph Classification

Model ReleasesDGX agent

arXiv:2605.02836v1 Announce Type: new Abstract: We introduce PLACE (Persistence-Landmark Analytic Classification Engine), a closed-form pipeline for classifying point clouds and graphs through their p

A decoupled diffusion planner that adapts to changing cost limits by using cost-conditioned generation for safety and reward gradients for performance

Model ReleasesDGX agent

arXiv:2605.02777v1 Announce Type: new Abstract: Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a s

A few days ago i wrote a post about how we use the LangGraph checkpointer to optimize storage. Today, LangGraph released v1.2 with a really …

Model ReleasesDGX agent

A few days ago i wrote a post about how we use the LangGraph checkpointer to optimize storage. Today, LangGraph released v1.2 with a really nice feature: DeltaChannel, a new channel type that stores o

A hybrid solution approach for the Integrated Healthcare Timetabling Competition 2024

Model ReleasesDGX agent

arXiv:2511.04685v2 Announce Type: replace Abstract: In this work, we present the solution approach for the Integrated Healthcare Timetabling Competition 2024 submitted by Team Twente, which ultimately

A Light Weight Multi-Features-View Convolution Neural Network For Plant Disease Identification

Model ReleasesDGX agent

arXiv:2605.00903v1 Announce Type: new Abstract: Agriculture is a key sector of the economies of developing countries. It serves as a primary source of income and employment for rural populations. Howe

A multilingual hallucination benchmark: MultiWikiQHalluA

Model ReleasesDGX agent

arXiv:2605.02504v1 Announce Type: new Abstract: Most hallucination evaluations focus on English, leaving it unclear whether findings transfer to lower-resource languages. We investigate faithfulness h

A Parameter-Free First-Order Algorithm for Non-Convex Optimization with ilde{mkern1mu O}(epsilon^{-5/3}) Global Rate

Model ReleasesDGX agent

arXiv:2605.02127v1 Announce Type: cross Abstract: We introduce PF-AGD, the first parameter-free, deterministic, accelerated first-order method to achieve O(epsilon^{-5/3}log(1/epsilon)) oracle complex

A Systematic Benchmark of Machine Transliteration Models for the Tajik-Farsi Language Pair: A Comparative Study from Rule-Based to Transformer Architectures

Model ReleasesDGX agent

arXiv:2605.02270v1 Announce Type: new Abstract: This paper presents the first comprehensive comparative analysis of modern machine learning architectures for transliteration between Tajik (Cyrillic sc

Accelerating battery research with an AI interface between FINALES and Kadi4Mat

Model ReleasesDGX agent

arXiv:2605.00909v1 Announce Type: cross Abstract: The time-consuming formation process critically impacts the longevity of sodium-ion coin cells and End Of Life (EOL) performance. This study aims to o

Accurate Legal Reasoning at Scale: Neuro-Symbolic Offloading and Structural Auditability for Robust Legal Adjudication

Model ReleasesDGX agent

arXiv:2605.02472v1 Announce Type: new Abstract: Legal texts often contain computational legal clauses--provisions whose understanding requires complex logic. While frontier Large Reasoning Models (LRM

Action Agent: Agentic Video Generation Meets Flow-Constrained Diffusion

Model ReleasesDGX agent

arXiv:2605.01477v1 Announce Type: new Abstract: We present Action Agent, a two-stage framework that unifies agentic navigation video generation with flow-constrained diffusion control for multi-embodi

Activation Compression in LLMs: Theoretical Analysis and Efficient Algorithm

Model ReleasesDGX agent

arXiv:2605.01255v1 Announce Type: new Abstract: Training large language models (LLMs) is highly memory-intensive, as training must store not only weights and optimizer states but also intermediate act

AdamO: A Collapse-Suppressed Optimizer for Offline RL

Model ReleasesDGX agent

arXiv:2605.01968v1 Announce Type: new Abstract: Offline reinforcement learning (RL) can fail spectacularly when bootstrapped temporal-difference (TD) updates amplify their own errors, driving the crit

Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis

Model ReleasesDGX agent

arXiv:2506.08849v4 Announce Type: replace Abstract: Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered

Adaptive Estimation and Inference in Semi-parametric Heterogeneous Clustered Multitask Learning via Neyman Orthogonality

Model ReleasesDGX agent

arXiv:2605.01907v1 Announce Type: cross Abstract: We study clustered multitask learning in a semiparametric setting where tasks share a latent cluster structure in their target parameters but exhibit

Adaptive Texture-aware Masking for Self-Supervised Learning in 3D Dental CBCT Analysis

Model ReleasesDGX agent

arXiv:2605.01741v1 Announce Type: new Abstract: Cone Beam Computed Tomography (CBCT) is pivotal for 3D diagnostic imaging in dentistry. However, the development of robust AI models for volumetric anal

Addressing Data Scarcity in Bangla Fake News Detection: An LLM-Based Dataset Augmentation Approach

Model ReleasesDGX agent

arXiv:2605.01292v1 Announce Type: new Abstract: The growing spread of misinformation in digital media highlights the need for reliable fake news detection systems, yet progress in under-resourced lang

Adoption and Use of LLMs at an Academic Medical Center

Model ReleasesDGX agent

arXiv:2602.00074v2 Announce Type: replace-cross Abstract: While large language models (LLMs) can support clinical documentation needs, standalone tools struggle with 'workflow friction' from manual da

AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.00425v1 Announce Type: new Abstract: Reinforcement learning (RL) has significantly advanced the ability of large language model (LLM) agents to interact with environments and solve multi-tu

Agent Factory Recap: How Gemma 4 Taught Itself Physics

Model ReleasesDGX agent

In this episode of The Agent Factory, Vlad Kolesnikov and I sat down with Omar Sanseviero from the Developer Experience team at Google DeepMind. We explored the groundbreaking release of Gemma 4: a ne

Agentic AI for Trip Planning Optimization Application

Model ReleasesDGX agent

arXiv:2605.00276v1 Announce Type: new Abstract: Trip planning for intelligent vehicles increasingly requires selecting optimal routes rather than merely producing feasible itineraries, as interacting

Agentopic: A Generative AI Agent Workflow for Explainable Topic Modeling

Model ReleasesDGX agent

arXiv:2605.00833v1 Announce Type: new Abstract: Agentopic is a novel agent-based workflow for explainable topic modeling that leverages the reasoning capabilities of Large Language Models (LLMs). Exis

AI agent startup Sierra valued at 15B in new 950M funding round

Model ReleasesDGX agent

Eight months after closing a 350 million funding round, Sierra Technologies Inc. today announced that it has raised an additional 950 million at a 15 billion valuation. Alphabet Inc.’s GV venture capi

AI-Driven Expansion and Application of the Alexandria Database

Model ReleasesDGX agent

arXiv:2512.09169v2 Announce Type: replace-cross Abstract: We present a novel multi-stage workflow for computational materials discovery that achieves a 99% success rate in identifying compounds within

AI Gateway lets you route to any model. On May 13 in SF, we're hosting a builder night powered by those models. Pick one, build, demo. Audie…

Model ReleasesDGX agent

AI Gateway lets you route to any model. On May 13 in SF, we're hosting a builder night powered by those models. Pick one, build, demo. Audience votes on best build. With @AnthropicAI, @MiniMax_AI, & @

Ai2 releases MolmoAct 2, enhancing robot intelligence in the real world

Model ReleasesDGX agent

Seattle-based artificial intelligence research institute Ai2, the Allen Institute for AI, today announced its next-generation open-source foundation artificial intelligence models, aimed at enabling r

Aligning LLMs with Biomedical Knowledge using Balanced Fine-Tuning

Model ReleasesDGX agent

arXiv:2511.21075v3 Announce Type: replace Abstract: Engineering LLMs to accelerate life sciences research requires a robust alignment with biomedical knowledge. We observe that biomedical text exhibit

An ALE-Consistent Graph Neural Operator-Transformer Framework for Fluid-Structure Interaction

Model ReleasesDGX agent

arXiv:2605.00937v1 Announce Type: cross Abstract: We propose an arbitrary Lagrangian-Eulerian (ALE)-consistent machine learning framework for long-term fluid-structure interaction (FSI) prediction on

An Efficient Metric for Data Quality Measurement in Imitation Learning

Model ReleasesDGX agent

arXiv:2605.01544v1 Announce Type: new Abstract: Imitation learning (IL) has seen remarkable progress, yet field deployment of IL-powered robots remains hindered by the challenge of out-of-distribution

AnchorD: Metric Grounding of Monocular Depth Using Factor Graphs

Model ReleasesDGX agent

arXiv:2605.02667v1 Announce Type: cross Abstract: Dense and accurate depth estimation is essential for robotic manipulation, grasping, and navigation, yet currently available depth sensors are prone t

ARA: Agentic Reproducibility Assessment For Scalable Support Of Scientific Peer-Review

Model ReleasesDGX agent

arXiv:2605.02651v1 Announce Type: cross Abstract: Scientific peer review increasingly struggles to assess reproducibility at the scale and complexity of modern research output. Evaluating reproducibil

ARIS: Agentic and Relationship Intelligence System for Social Robots

Model ReleasesDGX agent

arXiv:2605.00943v1 Announce Type: new Abstract: Foundational models have advanced social robotics, enabling richer perception and communicative interaction with users. However, current systems still s

← Previous
1…284285286287288…376
Next →