AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,611 results
9 Jun 2026

TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models

Model ReleasesDGX agent

arXiv:2510.27544v2 Announce Type: replace Abstract: Temporal reasoning involves understanding how systems evolve over time through input-driven state transitions. A key aspect is temporal causal reaso

The AI Epistemic Deference Index: A Continuous Measure of Sycophancy

Model ReleasesDGX agent

arXiv:2606.07897v1 Announce Type: new Abstract: Current AI models frequently exhibit epistemic sycophancy, endorsing claims to agree with a user. Existing evaluations typically measure this either by

The Injection Paradox: Brand-Level Suppression in Safety-Trained LLM Recommendations via RAG Context Injection

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.09204v1 Announce Type: new Abstract: We present a reproducible failure mode of safety training in RAG-based LLM recommendation -- the Injection Paradox -- in which prompt injections embedde

The Last Visible Pixel: Probing Fine-Scale Perception in Vision-Language Models

Model ReleasesDGX agent

arXiv:2606.07861v1 Announce Type: cross Abstract: Recent vision-language models (VLMs) excel at multimodal understanding and reasoning, yet their fine-grained visual perception remains underexplored.

The Montparnasse Algorithm for RNA Design

Model ReleasesDGX agent

arXiv:2606.07562v1 Announce Type: cross Abstract: RNA design consists of discovering a nucleotide sequence that optimizes predefined criteria, such as secondary structure. It is useful for synthetic b

The Sample Complexity of Parameter-Free Stochastic Convex Optimization

Model ReleasesDGX agent

arXiv:2506.11336v2 Announce Type: replace Abstract: We study the sample complexity of stochastic convex optimization when problem parameters such as the distance to optimality and the Lipschitz consta

TheoremBench: Evaluating LLMs on Theorem Proving in Formal Mathematics

Model ReleasesDGX agent

arXiv:2606.09450v1 Announce Type: new Abstract: LLMs have recently achieved strong results on formal proving benchmarks. However, existing evaluations remain heavily concentrated on competition-style

Theoretical Foundations of Continual Learning via Drift-Plus-Penalty

Model ReleasesDGX agent

arXiv:2606.08452v1 Announce Type: new Abstract: In many real-world settings, data streams are nonstationary and arrive sequentially, requiring learning systems to adapt continuously without retraining

There has been a lot of hand wringing on the appropriate valuation of SpaceX. Some large institutions believe SpaceX can only be valued at h…

Model ReleasesDGX agent

There has been a lot of hand wringing on the appropriate valuation of SpaceX. Some large institutions believe SpaceX can only be valued at half what the market seems to be willing to pay for it. Other

They ruled Iryna’s killer is incompetent to stand trial. The same system ruled this man was plenty competent enough to be released back into…

Model ReleasesDGX agent

They ruled Iryna’s killer is incompetent to stand trial. The same system ruled this man was plenty competent enough to be released back into society dozens of times. It’s past time to remove these lef

Thinking-Based Non-Thinking: Solving the Reward Hacking Problem in Training Hybrid Reasoning Models via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2601.04805v2 Announce Type: replace Abstract: Large reasoning models (LRMs) have attracted much attention due to their exceptional performance. However, their performance mainly stems from think

This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great…

Model ReleasesDGX agent

This is a super exciting release - Claude Fable 5 is the same underlying model as Mythos but with added safeguards. The benchmarks are great and it's SOTA on everything by a margin but I'll add that *

Tiger Data launches PostgreSQL extension designed for AI agents

Model ReleasesDGX agent

Tiger Data today introduced a managed PostgreSQL database service designed specifically for AI agents, saying conventional database architectures are poorly suited to a future in which software is inc

Today, we released Gemini 3.5 Live Translate, our latest audio model for live speech-to-speech translation. It supports over 70 languages an…

Model ReleasesDGX agent

Today, we released Gemini 3.5 Live Translate, our latest audio model for live speech-to-speech translation. It supports over 70 languages and starts translating as soon as you start talking, streaming

Token Sample Complexity of Attention

Model ReleasesDGX agent

arXiv:2512.10656v3 Announce Type: replace Abstract: As context windows in large language models continue to expand, it is essential to characterize how attention behaves at extreme sequence lengths. W

Towards Long-Horizon Vessel Trajectory and Destination Forecasting with Reasoning Large Language Models

Model ReleasesDGX agent

arXiv:2606.08633v1 Announce Type: new Abstract: Long-horizon maritime trajectory prediction is important for shipping management, logistics planning, and maritime risk analysis, yet month-level foreca

Towards Personalized Bangla Book Recommendation: A Large-Scale Heterogeneous Book Graph Dataset

Model ReleasesDGX agent

arXiv:2602.12129v2 Announce Type: replace-cross Abstract: Personalized book recommendation in Bangla literature has been constrained by the lack of structured, large-scale, and publicly available data

TQA-Bench: Evaluating LLMs for Multi-Table Question Answering

Model ReleasesDGX agent

arXiv:2411.19504v2 Announce Type: replace Abstract: The advance of large language models (LLMs) has unlocked great opportunities in complex multi-modal data management tasks, particularly in question

Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration

Model ReleasesDGX agent

arXiv:2606.09474v1 Announce Type: new Abstract: Generalized Few-Shot Semantic Segmentation (GFSS) has traditionally been approached as a representation-learning problem, requiring task-specific adapta

Trajectory-Refined Distillation

Model ReleasesDGX agent

arXiv:2606.08432v1 Announce Type: new Abstract: On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision alo

TriHead-GAN: A Generative Adversarial Network with Triple-Head Discriminator for Carbon Emission Time Series Generation

Model ReleasesDGX agent

arXiv:2606.07569v1 Announce Type: new Abstract: Accurate carbon emission monitoring is critical for climate policy and emerging regulatory mechanisms such as the EU Carbon Border Adjustment Mechanism,

TRL-Bench: Standardizing Cross-Paradigm Representation-Level Evaluation of Tabular Encoders

Model ReleasesDGX agent

arXiv:2606.09323v1 Announce Type: new Abstract: Tabular encoders are usually evaluated inside task-specific end-to-end pipelines, so models from different training paradigms are difficult to compare d

TT-DAC-PS: Twin-Target Deterministic Actor-Critic with Policy Smoothing for Optimal Trade Execution

Model ReleasesDGX agent

arXiv:2606.08379v1 Announce Type: new Abstract: This study addresses the optimal execution of large stock sell programs by introducing TT-DAC-PS (Twin-Target Deterministic Actor-Critic with Policy Smo

Understanding Benchmark Language Under Weakened Formal Semantics

Model ReleasesDGX agent

arXiv:2509.17455v2 Announce Type: replace-cross Abstract: State-of-the-art NLP benchmarks require interpretation of natural language that specifies conditions, procedures, and exceptions, often relyin

Understanding the Parameter Space Geometry of Transformers Encoding Boolean Functions

Model ReleasesDGX agent

arXiv:2606.08768v1 Announce Type: new Abstract: Transformers consistently fail to learn certain simple functions that are provably expressible with specific parameter settings. This gap between learna

Unification of Closed-Open Industrial Detection Scenarios: New Large-Scale Benchmarks,Challenges and Baselines

Model ReleasesDGX agent

arXiv:2606.07953v1 Announce Type: new Abstract: Large-scale Visual-Language Models (LVLMs) have achieved remarkable success in natural visual tasks, yet their application to industrial defect detectio

UniQL: Towards Dialect-Universal Benchmarking for Text-to-SQL

Model ReleasesDGX agent

arXiv:2606.08018v1 Announce Type: new Abstract: Existing text-to-SQL benchmarks are largely centered on SQLite, making it difficult to evaluate whether models can generalize across heterogeneous SQL d

Unsupervised Partner Design Enables Robust Ad-hoc Teamwork

Model ReleasesDGX agent

arXiv:2508.06336v2 Announce Type: replace-cross Abstract: We introduce Unsupervised Partner Design (UPD), a population-free multi-agent reinforcement learning method for robust ad-hoc teamwork. UPD ge

VATS: Exploiting Implicit Authority in Error-Path Injection via Systematic Mutation

Model ReleasesDGX agent

arXiv:2606.07992v1 Announce Type: new Abstract: As the Model Context Protocol (MCP) standardizes tool-calling for autonomous agents, it introduces a critical, unexamined attack surface: the error-hand

Version of AI tool too powerful for public released to public https://bbc.in/4xfGSlq

Model ReleasesDGX agent

A version of an AI tool deemed too powerful for public release was inadvertently made available to the public, according to a report from BBC News. The incident highlights concerns about controlling a

VideoWeaver: Evaluating and Evolving Skills for Agentic Long Video Generation

Model ReleasesDGX agent

arXiv:2606.08091v1 Announce Type: new Abstract: Recent agent frameworks such as Claude Code, Codex, and OpenClaw are strong at tool use and orchestration, but whether they can handle long video genera

Visual Template Inference for Data Extraction from Documents

Model ReleasesDGX agent

arXiv:2501.06659v2 Announce Type: replace-cross Abstract: Many templatized documents are programmatically generated from structured data following a visual template. Such documents include invoices, t

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?

Model ReleasesDGX agent

arXiv:2606.07872v1 Announce Type: new Abstract: When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual e

VisualLeakBench: Reproducible Action-Boundary Propagation Failures in Vision-Language Agents

Model ReleasesDGX agent

arXiv:2606.07595v1 Announce Type: cross Abstract: Vision-language agents increasingly consume screenshots, documents, and user interfaces before writing to memory, sending messages, or invoking extern

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.08094v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware

VoLo: A Physical Orchestrator for Open-Vocabulary Long-Horizon Manipulation

Model ReleasesDGX agent

arXiv:2606.07723v1 Announce Type: new Abstract: Open-vocabulary long-horizon manipulation requires robots to reason over flexible instructions and complex multi-object scenes while adaptively planning

We encourage developers to share their builds with us and give feedback to shape future iterations. Let’s shape the future of sovereign AI t…

Model ReleasesDGX agent

We encourage developers to share their builds with us and give feedback to shape future iterations. Let’s shape the future of sovereign AI together. Download: https://huggingface.co/CohereLabs/North-M

We talk a lot about how important it is to set up self-verification loops. Especially in the age of powerful models that can run for long pe…

Model ReleasesDGX agent

We talk a lot about how important it is to set up self-verification loops. Especially in the age of powerful models that can run for long periods of time, self-verification is a key ingredient that en

WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces

Model ReleasesDGX agent

arXiv:2606.09426v1 Announce Type: new Abstract: Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and ext

We're getting Fable 5 before GTA 6

Model ReleasesDGX agent

This post humorously suggests that Fable 5 will release before Grand Theft Auto 6, likely commenting on the extended development timelines of both highly anticipated games. The statement reflects comm

We're hosting Claude Fable 5 Build Day in San Francisco on June 13. Point Fable 5 at a problem worth solving and build a solution with Claud…

Model ReleasesDGX agent

We're hosting Claude Fable 5 Build Day in San Francisco on June 13. Point Fable 5 at a problem worth solving and build a solution with Claude Code. The Anthropic team will be in the room, with a chanc

We've got 13 days to burn as much tokens as humanly possible on Claude Max plans Before they revert to API based billing 💀

Model ReleasesDGX agent

We've got 13 days to burn as much tokens as humanly possible on Claude Max plans Before they revert to API based billing 💀 Introducing Claude Fable 5: a Mythos-class model that we’ve made safe for gen

What Codex unlocks for Notion

Model ReleasesDGX agent

OpenAI's Codex model enables Notion to add AI-powered capabilities to its workspace platform, allowing users to automate tasks and generate content through natural language commands. This integration

What it feels like to work with Mythos

Model ReleasesDGX agent

This article by Ethan Mollick describes the user experience and practical workflow of working with Mythos, an AI system. It likely covers the system's capabilities, interface, strengths, limitations,

When Behavioral Safety Evaluation Fails: A Representation-Level Perspective

Model ReleasesDGX agent

arXiv:2606.08044v1 Announce Type: cross Abstract: Large Language Model (LLM) safety has often been evaluated at the behavior level, which provides limited evidence of internal robustness, as these eva

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

Model ReleasesDGX agent

arXiv:2602.08235v2 Announce Type: replace-cross Abstract: Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unin

When Do Local Score Models Extrapolate Across Size? A Diagnostic Theory and Benchmark

Model ReleasesDGX agent

arXiv:2606.09705v1 Announce Type: new Abstract: Scientific generative modeling often requires size transfer, where models trained on small systems are evaluated on larger ones. While translation-invar

Where Does the Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving

Model ReleasesDGX agent

arXiv:2606.09644v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) achieve strong results on visual reasoning benchmarks, but answer accuracy alone does not indicate whether a

Where Instruction Hierarchy Breaks: Diagnosing and Repairing Failures in Reasoning Language Models

Model ReleasesDGX agent

arXiv:2606.07808v1 Announce Type: new Abstract: Reasoning language models deployed in agentic workflows must follow an instruction hierarchy: when instructions from different sources conflict, the mod

Who Earns the Safety? Intervention-Aware Quantum Predictive Control with Safety Attribution

Model ReleasesDGX agent

arXiv:2606.09778v1 Announce Type: cross Abstract: Hard safety filters are increasingly placed downstream of learned controllers to guarantee constraint satisfaction at run time. Yet a filtered control

Wordle 1,815 6/6 ⬛⬛🟨⬛⬛ ⬛🟨⬛⬛⬛ ⬛🟩⬛🟩⬛ 🟩🟩⬛🟩⬛ 🟩🟩🟩🟩⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post shows a completed Wordle game (#1,815) solved on the sixth and final attempt, with the emoji grid indicating which letters were correct, misplaced, or absent in each guess. Anthropic shared

XCR-Bench: Benchmarking Cross-Cultural Reasoning in LLMs via Culture-Specific Items and Hall's Triad

Model ReleasesDGX agent

arXiv:2601.14063v2 Announce Type: replace-cross Abstract: Cross-cultural competence in large language models (LLMs) requires understanding and adapting Culture-Specific Items (CSIs) across varying cul

You can try Claude Fable 5 as part of Devin Cloud’s Ultra agent. Devin Ultra is our smartest and most capable agent, which excels at long-ho…

Model ReleasesDGX agent

You can try Claude Fable 5 as part of Devin Cloud’s Ultra agent. Devin Ultra is our smartest and most capable agent, which excels at long-horizon tasks and debugging. We tuned the harness so Ultra cos

Your Model Already Knows: Attention-Guided Safety Filter for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.09749v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated impressive end-to-end performance across a variety of robotic manipulation tasks. However, these

Zero-Parameter Geometric Gating for Temporally Stable Low-Altitude UAV Video Semantic Segmentation

Model ReleasesDGX agent

arXiv:2606.09162v1 Announce Type: new Abstract: Video semantic segmentation for low-altitude UAVs requires temporal consistency, yet dense optical flow introduces spatially structured noise in the pla

Zero-Shot Learning in Industrial Scenarios: New Large-Scale Benchmark, Challenges and Baseline

Model ReleasesDGX agent

arXiv:2606.07965v1 Announce Type: new Abstract: Large Visual Language Models (LVLMs) have achieved remarkable success in vision tasks. However, the significant differences between industrial and natur

Zero-Shot Semantic Re-Identification for Autonomous Driving: A VLM Baseline Study

Model ReleasesDGX agent

arXiv:2606.09362v1 Announce Type: new Abstract: Re-Identification (ReID) in autonomous driving is typically formulated as a visual matching problem, where observations of vehicles, pedestrians, and cy

ZIPP:Zero-shot Image Personalization from Personas

Model ReleasesDGX agent

arXiv:2606.08841v1 Announce Type: new Abstract: Text-to-image diffusion models are increasingly deployed in open-ended creative contexts, yet their outputs remain impersonal, optimized for aggregate a

Zscaler launches AI Broker and Endpoint AI Security for agents

Model ReleasesDGX agent

Zscaler Inc. today unveiled a set of products designed to secure autonomous artificial intelligence agents, with the cybersecurity company claiming it has built the industry’s first complete zero-trus

8 Jun 2026

3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing

Model ReleasesDGX agent

arXiv:2606.07115v1 Announce Type: new Abstract: Despite recent progress in 3D generation, intuitive editing of existing shapes remains limited. Unlike images, which benefit from well-established inpai

← Previous
1…160161162163164…377
Next →