AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,620 results
29 May 2026

Text-Preserving Lossy Text Compression: A Study of Strategic Deletion and LLM Reconstruction

Model ReleasesDGX agent

arXiv:2605.29000v1 Announce Type: new Abstract: Traditional lossless text compression preserves every byte, but its gains on natural language are often modest in realistic operating regimes. We study

The blast bent those steel beams on the tower inwards:

Model ReleasesDGX agent

The blast bent those steel beams on the tower inwards: First look at LC-36 from the air this morning after the explosion of New Glenn last night during a failed hotfire test. Visible is the wreckage f

The Chain Holds, the Answer Folds: Trace-Answer Dissociation in Reasoning Models Under Adversarial Pressure

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.29087v1 Announce Type: new Abstract: Reasoning models are evaluated on single-turn benchmarks but deployed in multi-turn dialogue, where users push back on correct answers. Under sustained

The Cognitive Categorical Transformer: Category-Theoretic Inductive Biases for Language Modeling

Model ReleasesDGX agent

arXiv:2605.28864v1 Announce Type: new Abstract: The Cognitive Categorical Transformer (CCT) is a 306M-parameter architecture that augments a pretrained GPT-2 Small backbone with cognitively grounded c

The Curse of Helpfulness: Inverse Scaling Law in Robustness to Distractor Instructions via DistractionIF

Model ReleasesDGX agent

arXiv:2605.29491v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in agentic and retrieval-augmented generation (RAG) systems, where they must execute user-specifi

The Good, the Bad, and the Ugly of Markov Boundary for Tabular Prediction

Model ReleasesDGX agent

arXiv:2605.29411v1 Announce Type: cross Abstract: Under standard graphical assumptions, the Markov boundary of a target variable is the smallest set of features that renders every other feature redund

The Hamilton-Jacobi Theory of Deep Learning

Model ReleasesDGX agent

arXiv:2605.28983v1 Announce Type: cross Abstract: In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selec

The improvements run wide. Across all major European languages, Command A+ consistently pulls ahead of competitors on WMT24++ (xCOMET-XL): …

Model ReleasesDGX agent

The improvements run wide. Across all major European languages, Command A+ consistently pulls ahead of competitors on WMT24++ (xCOMET-XL): 🇫🇷 +2.4 pts in French 🇪🇸 +1.9 pts in Spanish 🇩🇪 +0.9 pts in G

The Open Motion Planning Library 2.0

Model ReleasesDGX agent

arXiv:2605.29301v1 Announce Type: new Abstract: The Open Motion Planning Library (OMPL), first released in 2008, has become a cornerstone of the motion planning community, providing implementations of

The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More

Model ReleasesDGX agent

arXiv:2603.23971v2 Announce Type: replace-cross Abstract: Developers and consumers increasingly choose reasoning models (RMs) based on their listed API prices. However, how accurately do these prices

The story gets bigger beyond Europe. Command A+ makes major gains in high-impact non-Latin languages; outperforming Mistral Medium 3.5 in Ko…

Model ReleasesDGX agent

The story gets bigger beyond Europe. Command A+ makes major gains in high-impact non-Latin languages; outperforming Mistral Medium 3.5 in Korean, Japanese, Hebrew, Chinese, and Arabic. For Arabic, tha

The team at @llama_index built an awesome template using LlamaParse and the new Managed Agents in the Gemini API. See how they built an agen…

Model ReleasesDGX agent

The team at @llama_index built an awesome template using LlamaParse and the new Managed Agents in the Gemini API. See how they built an agent that can tackle unstructured documents. 📄↓ 🚀 The team at @

The Trust Paradox: How CS Researchers Engage LLM Leaderboards

Model ReleasesDGX agent

arXiv:2605.28966v1 Announce Type: new Abstract: Large language model (LLM) leaderboards rank AI models using standardized benchmarks and have become highly visible across computer science, despite kno

The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

Model ReleasesDGX agent

arXiv:2602.15382v2 Announce Type: replace Abstract: Multi-Agent Systems (MAS) powered by Large Language Models have unlocked advanced collaborative reasoning, yet they remain bottlenecked by discrete

This is a diary entry to myself, so I remember what AI was like today. It's just going to be a bullet-list stream of consciousness. - There …

Model ReleasesDGX agent

This is a diary entry to myself, so I remember what AI was like today. It's just going to be a bullet-list stream of consciousness. - There are still so many leaders that have never seen an agent run

this was my pypi hardening strategy for @activegraphai: - spin up a new @replit - point at docs page, ask it to build something - ask it to …

Model ReleasesDGX agent

this was my pypi hardening strategy for @activegraphai: - spin up a new @replit - point at docs page, ask it to build something - ask it to write a feedback report to package builder - feed that feedb

Three-dimensional Conditional Diffusion Models for Cosmological 21 cm Lightcone Emulation

Model ReleasesDGX agent

arXiv:2605.29016v1 Announce Type: cross Abstract: We investigate conditional diffusion modeling for three-dimensional 21 cm lightcone emulation, focusing on cubes with a sky-plane size of 64imes64 and

TIMEGATE: Sustainable Time-Boxed Promotion Gates for Continual ML Adaptation Under Resource Constraints

Model ReleasesDGX agent

arXiv:2605.29183v1 Announce Type: cross Abstract: As machine learning(ML) systems evolve to continual adaptation, each re-training cycle uses compute, annotation, and energy. We introduce TIMEGATE, a

Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection

Model ReleasesDGX agent

arXiv:2605.30344v1 Announce Type: new Abstract: Recent advances in Vision-Language Models (VLMs) have achieved impressive performance across many tasks, yet prior studies report unsatisfactory perform

Token-Level Generalization in LoRA Adapter Backdoors: Attack Characterization and Behavioral Detection

Model ReleasesDGX agent

arXiv:2605.30189v1 Announce Type: cross Abstract: We show that LoRA adapters, the dominant distribution format for fine-tuned LLMs, can be reliably backdoored through training data poisoning while pre

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but i…

Model ReleasesDGX agent

Took me a while to figure out what all the ESMFold2 rage was about. At first, the benchmarking data didn't look super remarkable to me but it turns there are many impressive aspects: - Fully open sour

Toward Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

Model ReleasesDGX agent

arXiv:2605.29230v1 Announce Type: cross Abstract: Age estimation from facial images typically relies on training data that includes images of minors, a practice that raises serious ethical, legal, and

TRACE: Toulmin-based Reasoning Assessment through Constructive Elements for LLM CoT Evaluation

Model ReleasesDGX agent

arXiv:2605.29656v1 Announce Type: new Abstract: Evaluating open-ended outputs from large language models (LLMs) remains challenging due to the absence of ground truth. Existing metrics rely on final-a

Training Deliberative Monitors for Black-Box Scheming Detection

Model ReleasesDGX agent

arXiv:2605.29601v1 Announce Type: cross Abstract: As autonomous agents become more capable of performing real-world tasks, distinguishing scheming behavior from benign task pursuit may become a centra

Trends in AI and Human-AI Interaction in Clinical Trials -- A Hybrid Human-AI Exploration

Model ReleasesDGX agent

arXiv:2605.29096v1 Announce Type: new Abstract: This paper examines records retrieved from the ClinicalTrials.gov registry to characterize temporal trends in AI terminology and the geographical distri

UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning

Model ReleasesDGX agent

arXiv:2605.29170v1 Announce Type: cross Abstract: Legal NLP benchmarks are overwhelmingly English-centric, leaving failure modes in morphologically rich, non-Latin-script languages undetected. We intr

Uncertainty-Aware Transfer Learning for Cross-Building Energy Forecasting: Toward Robust and Scalable District-Level Energy Management

Model ReleasesDGX agent

arXiv:2605.29733v1 Announce Type: new Abstract: Scaling data-driven energy forecasting to district level requires models that can be re-used across buildings with minimal target-domain data and honest

Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs

Model ReleasesDGX agent

arXiv:2605.29708v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) LLMs rely on sparse, router-driven expert activation, yet how safety alignment interacts with routed expert specialization rema

VE2VF: Vision-Enabled to Vision-Free Distillation via Real-world Reinforcement Learning for Robust Contact-Rich Manipulation

Model ReleasesDGX agent

arXiv:2605.29564v1 Announce Type: new Abstract: When using reinforcement learning (RL) for contact-rich robotic manipulation, vision can provide task-relevant information that accelerates learning bey

Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering

Model ReleasesDGX agent

arXiv:2605.29648v1 Announce Type: new Abstract: Applying reinforcement learning to improve factual accuracy in knowledge-intensive question answering faces a reward design dilemma. Response-level rewa

VIDEO: @BlueOrigin major New Glenn static fire anomaly at Launch Complex-36 📷 @JerryPikePhoto/@NASASpaceflight

Model ReleasesDGX agent

VIDEO: @BlueOrigin major New Glenn static fire anomaly at Launch Complex-36 📷 @JerryPikePhoto/@NASASpaceflight Media ANOMALY: @BlueOrigin have suffered a CATASTROPHIC EXPLOSION AT LAUNCH COMPLEX-36 📷

Video Individual Counting and Tracking from Moving Drones: A Benchmark and Methods

Model ReleasesDGX agent

arXiv:2601.12500v2 Announce Type: replace Abstract: Counting and tracking dense crowds in large-scale scenes is a highly practical yet challenging problem. Existing methods mostly rely on fixed-camera

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

Model ReleasesDGX agent

arXiv:2605.30256v1 Announce Type: cross Abstract: Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonve

VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring over Wearable Health Data

Model ReleasesDGX agent

arXiv:2605.29483v1 Announce Type: new Abstract: Wearable devices enable continuous monitoring of physiological signals such as ECG and PPG, but existing mHealth systems are largely limited to task-spe

VLAConf: Calibrated Task-Success Confidence for Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2605.29605v1 Announce Type: new Abstract: Confidence estimation for Vision-Language-Action (VLA) models is essential for robots to perform manipulation tasks in the open world, providing crucial

VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation

Model ReleasesDGX agent

arXiv:2605.30317v1 Announce Type: new Abstract: Autoregressive image and video generators are trained with teacher-forced histories but must sample from their own generated prefixes at inference time,

WASHH: An Anchor-Aware Whale-Guided Selection Hyper-Heuristic for Continuous Optimization and SVC Configuration

Model ReleasesDGX agent

arXiv:2605.28844v1 Announce Type: cross Abstract: Learning-assisted algorithm design often has to make reliable search decisions under small evaluation budgets, where committing to a single metaheuris

Watch your avatar speak Spanish, English and Japanese https://x.com/DotCSV/status/2059676610400231599?s=20

Model ReleasesDGX agent

Watch your avatar speak Spanish, English and Japanese https://x.com/DotCSV/status/2059676610400231599?s=20 Sigo jugando con Omni! Efectivamente el modelo desbloquea un montón de casos de uso (e.g. tra

We’re taking steps to accelerate defensive progress in biology: - Launching Rosalind Biodefense to help trusted builders develop new biodefe…

Model ReleasesDGX agent

We’re taking steps to accelerate defensive progress in biology: - Launching Rosalind Biodefense to help trusted builders develop new biodefense and pandemic preparedness capabilities. - Expanding trus

What drives performance in molecular MPNNs? An operator-level factorial benchmark

Model ReleasesDGX agent

arXiv:2605.30195v1 Announce Type: cross Abstract: Message-passing neural networks (MPNNs) are widely used for molecular property prediction, but their deployment as monolithic architectures makes it d

When LLM Reward Design Fails: Diagnostic-Driven Refinement for Sparse Structured RL

Model ReleasesDGX agent

arXiv:2605.28918v1 Announce Type: new Abstract: For sparse, structured reinforcement-learning tasks with semantic reward-function interfaces, LLM-generated reward shaping is better framed as debugging

When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making

Model ReleasesDGX agent

arXiv:2603.16673v4 Announce Type: replace-cross Abstract: Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models

Model ReleasesDGX agent

arXiv:2605.30219v1 Announce Type: new Abstract: Long-horizon interactions require language models to manage accumulating information: when to update their state, when to preserve their state, and what

When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models

Model ReleasesDGX agent

arXiv:2601.00065v3 Announce Type: replace-cross Abstract: Tokenizer transplant in cross-vocabulary model composition reconstructs donor-only embedding rows as weighted combinations over shared lexical

When we say “LiteParse runs everywhere,” we mean it. Our WASM package is lightweight, minimal, and built for browser and edge runtimes, whic…

Model ReleasesDGX agent

When we say “LiteParse runs everywhere,” we mean it. Our WASM package is lightweight, minimal, and built for browser and edge runtimes, which makes it a perfect fit for @cloudflare Workers. Using WebA

When you are talking to an LLM, you are speaking to a synthesized work of interactive fiction, not a real being.

Model ReleasesDGX agent

When you are talking to an LLM, you are speaking to a synthesized work of interactive fiction, not a real being. ChatGPT, Claude, and Sydney are not their neural networks. If any LLM claims to be cons

Who can we trust? LLM-as-a-jury for Comparative Assessment

Model ReleasesDGX agent

arXiv:2602.16610v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly applied as automatic evaluators for natural language generation assessment often using pairwise

Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.30161v1 Announce Type: new Abstract: Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D und

Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

Model ReleasesDGX agent

arXiv:2605.29744v1 Announce Type: new Abstract: The impressive performance of generalist large language models (LLMs) such as GPT and Claude in healthcare raises a critical question: will domain-speci

Windows users, this one’s for you. Computer use now works on Windows, so Codex can take action on your Windows computer. And with Windows su…

Model ReleasesDGX agent

Windows users, this one’s for you. Computer use now works on Windows, so Codex can take action on your Windows computer. And with Windows support for Codex in the ChatGPT mobile app, you can start, re

Wordle 1,804 4/6 ⬛🟨⬛⬛⬛ ⬛🟩⬛⬛🟨 ⬛🟩🟩🟩🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This entry documents a Wordle game result where Anthropic solved puzzle #1,804 in 4 attempts, using color-coded feedback (⬛ = incorrect letter, 🟨 = correct letter wrong position, 🟩 = correct letter co

World Models in Words: Auditing Physical State-Transition Commitments in Vision-Language Models

Model ReleasesDGX agent

arXiv:2605.29585v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to answer questions about physical scenes, yet most evaluations reduce performance to a final answer

YoCausal: How Far is Video Generation from World Model? A Causality Perspective

Model ReleasesDGX agent

arXiv:2605.30346v1 Announce Type: new Abstract: As video diffusion models (VDMs) advance toward world models, a key question arises: do they truly understand causality, or merely overfit to statistica

28 May 2026

$65B private round More than double the size of the largest IPO ever

Model ReleasesDGX agent

65B private round More than double the size of the largest IPO ever We've raised 65 billion in Series H funding at a $965 billion post-money valuation, led by @AltimeterCap, Dragoneer, @Greenoaks, and

A Bayesian Nonparametric Perspective on Mahalanobis Distance for Out of Distribution Detection

Model ReleasesDGX agent

arXiv:2502.08695v2 Announce Type: replace-cross Abstract: Bayesian nonparametric methods are naturally suited to the problem of out-of-distribution (OOD) detection. However, these techniques have larg

A Broader View of Thompson Sampling

Model ReleasesDGX agent

arXiv:2510.07208v2 Announce Type: replace Abstract: Thompson Sampling is one of the most widely used and studied bandit algorithms, known for its simple structure, low regret performance, and solid th

A Fresh Look at Lamarckian Evolution and the Baldwin Effect

Model ReleasesDGX agent

arXiv:2605.28703v1 Announce Type: cross Abstract: Baldwinian and Lamarckian evolution have existed for a long time in evolutionary algorithms (EAs) without ever dominating the academic literature or p

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Model ReleasesDGX agent

arXiv:2605.28556v1 Announce Type: new Abstract: As agent capabilities advance, existing benchmarks, such as au^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remain

A Multi-dimensional Framework for Evaluating Generalization in EEG Foundation Models

Model ReleasesDGX agent

arXiv:2605.28563v1 Announce Type: cross Abstract: Evaluating foundation models under appropriate adaptation settings is essential for understanding the quality and transferability of the learned repre

A Query Engine for the Agents

Model ReleasesDGX agent

arXiv:2605.27785v1 Announce Type: new Abstract: The fastest-growing data in production today is unstructured text: agent traces, chat logs, reasoning chains, model outputs. People want to analyze it,

← Previous
1…197198199200201…377
Next →