AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,553 results
6 May 2026

ReCode: Reinforcing Code Generation with Reasoning-Process Rewards

Model ReleasesDGX agent

arXiv:2508.05170v3 Announce Type: replace-cross Abstract: In practice, rigorous reasoning is often a key driver of correct code, while Reinforcement Learning (RL) for code generation often neglects op

ReLeaf: Benchmarking Leaf Segmentation across Domains and Species

Model ReleasesDGX agent

arXiv:2605.03784v1 Announce Type: new Abstract: Rising global food demand and growing climate pressure increase the need for sustainable, precise agricultural practices. Automated, individualized plan

Reproducing Complex Set-Compositional Information Retrieval

Model ReleasesDGX agent

arXiv:2605.03824v1 Announce Type: new Abstract: Complex information needs may involve set-compositional queries using conjunction, disjunction, and exclusion, yet it remains unclear whether current re


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems

Model ReleasesDGX agent

arXiv:2605.04018v1 Announce Type: new Abstract: Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capabilit

Retrieval and Multi-Hop Reasoning in 1M-Token Context Windows: Evaluating LLMs on Classical Chinese Text

Model ReleasesDGX agent

arXiv:2605.02173v1 Announce Type: new Abstract: We evaluate the long-context retrieval and reasoning capabilities of five frontier large language models with advertised 1M-token context windows on a c

Retrieval-Augmented LLMs for Security Incident Analysis

Model ReleasesDGX agent

arXiv:2603.18196v3 Announce Type: replace-cross Abstract: Investigating cybersecurity incidents requires collecting and analyzing evidence from multiple log sources, including intrusion detection aler

Reward Hacking Benchmark: Measuring Exploits in LLM Agents with Tool Use

Model ReleasesDGX agent

arXiv:2605.02964v1 Announce Type: new Abstract: Reinforcement learning (RL) trained language model agents with tool access are increasingly deployed in coding assistants, research tools, and autonomou

RFPrompt: Prompt-Based Expert Adaptation of the Large Wireless Model for Modulation Classification

Model ReleasesDGX agent

arXiv:2605.03279v1 Announce Type: new Abstract: Automatic modulation classification (AMC) in real-world deployments demands robustness to distribution shifts arising from hardware impairments, unseen

RoboAlign-R1: Distilled Multimodal Reward Alignment for Robot Video World Models

Model ReleasesDGX agent

arXiv:2605.03821v1 Announce Type: new Abstract: Existing robot video world models are typically trained with low-level objectives such as reconstruction and perceptual similarity, which are poorly ali

RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation

Model ReleasesDGX agent

arXiv:2507.00435v2 Announce Type: replace-cross Abstract: We introduce RoboEval, a structured evaluation framework and benchmark for robotic manipulation that augments binary success with principled b

Robust Language Identification for Romansh Varieties

Model ReleasesDGX agent

arXiv:2603.15969v2 Announce Type: replace Abstract: The Romansh language has several regional varieties, called idioms, which sometimes have limited mutual intelligibility. Despite this linguistic div

S2O: Early Stopping for Sparse Attention via Online Permutation

Model ReleasesDGX agent

arXiv:2602.22575v2 Announce Type: replace Abstract: Attention scales quadratically with sequence length, fundamentally limiting long-context inference. Existing block-granularity sparsification can re

Safety and accuracy follow different scaling laws in clinical large language models

Model ReleasesDGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

SAM-NER: Semantic Archetype Mediation for Zero-Shot Named Entity Recognition

Model ReleasesDGX agent

arXiv:2605.03706v1 Announce Type: new Abstract: Zero-shot Named Entity Recognition (ZS-NER) remains brittle under domain and schema shifts, where unseen label definitions often misalign with a large l

Scaling Unsupervised Multi-Source Federated Domain Adaptation through Group-Wise Discrepancy Minimization

Model ReleasesDGX agent

arXiv:2510.08150v3 Announce Type: replace Abstract: Unsupervised multi-source domain adaptation (UMDA) leverages labeled data from multiple source domains to generalize to an unlabeled target. While f

SCGNN: Semantic Consistency enhanced Graph Neural Network Guided by Granular-ball Computing

Model ReleasesDGX agent

arXiv:2605.02617v2 Announce Type: new Abstract: Capturing semantic consistency among nodes is crucial for effective graph representation learning. Existing approaches typically rely on k-nearest neigh

See what these young builders are doing with AI now that everyone can just build things. https://chatgpt.com/futures/

Model ReleasesDGX agent

This OpenAI social media post highlights how young builders and developers are leveraging AI tools to create projects and applications now that AI development has become more accessible to non-experts

Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?

Model ReleasesDGX agent

arXiv:2605.00964v1 Announce Type: cross Abstract: Much research on LLMs has focused on increasing benchmark performance. However, the evaluation of such models in real-world collaborative human-AI wor

Self-Mined Hardness for Safety Fine-Tuning

Model ReleasesDGX agent

arXiv:2605.03226v1 Announce Type: new Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's diff

Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

Model ReleasesDGX agent

arXiv:2605.03189v1 Announce Type: new Abstract: Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While sev

Simulated Students in Tutoring Dialogues: Substance or Illusion?

Model ReleasesDGX agent

arXiv:2601.04025v2 Announce Type: replace Abstract: Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires

Singular Bank helps bankers move fast with ChatGPT and Codex

Model ReleasesDGX agent

Singular Bank leverages OpenAI's ChatGPT and Codex to accelerate banking operations and developer productivity. The integration enables bankers and software engineers to streamline workflows, automate

so cute!

Model ReleasesDGX agent

so cute! It's pretty magical how kids interact with tech. They just expect it to work for them. Here's mine talking to Reachy Mini via Gemini Live API. I didn't tell them how, they just did it. That's

Soft Tournament Equilibrium

Model ReleasesDGX agent

arXiv:2604.04328v3 Announce Type: replace-cross Abstract: The evaluation of general-purpose artificial agents, particularly those based on LLMs, presents a significant challenge due to the non-transit

SpaceXAI will provide @AnthropicAI with access to Colossus 1, one of the world’s largest and fastest-deployed AI supercomputers, to provide …

Model ReleasesDGX agent

SpaceXAI will provide @AnthropicAI with access to Colossus 1, one of the world’s largest and fastest-deployed AI supercomputers, to provide additional capacity for Claude → http://x.ai/news/anthropic-

Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning

Model ReleasesDGX agent

arXiv:2605.03229v1 Announce Type: new Abstract: Adapting a pretrained language model to a new task often hurts the general capabilities it already had, a problem known as catastrophic forgetting. Spar

Stable Multimodal Graph Unlearning via Feature-Dimension Aware Quantile Selection

Model ReleasesDGX agent

arXiv:2605.03303v1 Announce Type: new Abstract: Graph unlearning remains a critical technique for supporting privacy-preserving and sustainable multimodal graph learning. However, we observe that exis

StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing

Model ReleasesDGX agent

arXiv:2605.02904v1 Announce Type: new Abstract: We present StateSMix, a fully self-contained lossless compressor that couples an online-trained Mamba-style State Space Model (SSM) with sparse n-gram c

StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning

Model ReleasesDGX agent

arXiv:2605.03927v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural

Strong Opinions, Loosely Held on Agent + Harness Engineering: 1. You can outperform any default harness+model (including codex & claude code…

Model ReleasesDGX agent

Strong Opinions, Loosely Held on Agent + Harness Engineering: 1. You can outperform any default harness+model (including codex & claude code) on pretty much any Task by engineering the harness around

SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.03534v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, but retrieval is not verification: a passage can be topical and still fail t

@swyx @maheshmurag deep dive on the new dreaming feature in beta today https://x.com/claudeai/status/2052067400690851842?s=46

Model ReleasesDGX agent

@swyx @maheshmurag deep dive on the new dreaming feature in beta today https://x.com/claudeai/status/2052067400690851842?s=46 Dreaming reviews your agent's past sessions, extracts patterns, and curate

Take: stop searching for the “best model' or “best harness”. Model release → harness breaks → refind your optimal config. Most teams handle …

Model ReleasesDGX agent

Take: stop searching for the “best model' or “best harness”. Model release → harness breaks → refind your optimal config. Most teams handle this with manual local search, with poor visibility into acc

Task-Aware Scanning Parameter Configuration for Robotic Inspection Using Vision Language Embeddings and Hyperdimensional Computing

Model ReleasesDGX agent

arXiv:2605.03909v1 Announce Type: cross Abstract: Robotic laser profiling is widely used for dimensional verification and surface inspection, yet measurement fidelity is often dominated by sensor conf

Tenability and Weak Semantics: Modeling Non-uniform Defense -- Extended Version

Model ReleasesDGX agent

arXiv:2605.02024v1 Announce Type: new Abstract: In Dung-style abstract argumentation, various semantics capture notions of acceptability of arguments. The admissibility semantics capture the notion th

The Blueprint: Translating stream-of-conscious speech into responsive, actionable task lists

Model ReleasesDGX agent

Welcome to The Blueprint, a new feature where we highlight how Google Cloud customers are tackling unique and common challenges across industries using the latest AI and cloud technologies. We hope to

The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate

Model ReleasesDGX agent

arXiv:2605.00914v1 Announce Type: cross Abstract: Multi-agent debate, where teams of LLMs iteratively exchange rationales and vote on answers, is widely deployed under the assumption that peer review

The Dynamic Gist-Based Memory Model (DGMM): A Memory-Centric Architecture for Artificial Intelligence

Model ReleasesDGX agent

arXiv:2605.02106v1 Announce Type: new Abstract: Contemporary artificial intelligence systems achieve strong performance through large-scale parameterization, retrieval augmentation, and training on ex

The full Structured Output Benchmark dataset is now on @huggingface https://huggingface.co/datasets/interfaze-ai/sob

Model ReleasesDGX agent

The Structured Output Benchmark (SOB) dataset has been released and made available on Hugging Face. This dataset, hosted by Interfaze AI, likely provides benchmark data for evaluating models' ability

The last sentence in this abstract is really important, in a way that professional programmers will immediately recognize: the models favore…

Model ReleasesDGX agent

The last sentence in this abstract is really important, in a way that professional programmers will immediately recognize: the models favored big single files rather than breaking things into modules.

The Oracle's Fingerprint: Correlated AI Forecasting Errors and the Limits of Bias Transmission

Model ReleasesDGX agent

arXiv:2605.00844v1 Announce Type: cross Abstract: When large language models (LLMs) are consulted as forecasting tools, the independence of individual errors -- the foundation of collective intelligen

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development

Model ReleasesDGX agent

arXiv:2605.01160v1 Announce Type: cross Abstract: Since 2022, AI-powered coding assistants have produced contradictory evidence: controlled studies report 20-56% productivity gains on well-scoped task

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It

Model ReleasesDGX agent

arXiv:2605.03258v1 Announce Type: cross Abstract: Large language models often fail at simple counting tasks, even when the items to count are explicitly present in the prompt. We investigate whether t

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

Model ReleasesDGX agent

arXiv:2605.03073v1 Announce Type: new Abstract: Niche-domain Indic ASR -- digit strings, currency amounts, addresses, brand names, English/Indic codemix -- is under-served by both open-source SOTA and

this deepagents deploy https://docs.langchain.com/oss/python/deepagents/deploy (or at least directionally where we want to take it) what's m…

Model ReleasesDGX agent

this deepagents deploy https://docs.langchain.com/oss/python/deepagents/deploy (or at least directionally where we want to take it) what's missing? give us feedback! can someone PLEASE launch OS claud

This week made something clear: you shouldn't take what most tech ceos are saying publicly seriously!

Model ReleasesDGX agent

This week made something clear: you shouldn't take what most tech ceos are saying publicly seriously! From “Anthropic is Misanthropic” to “Claude is good for humanity and was impressed.” Most ironic o

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

Model ReleasesDGX agent

arXiv:2605.01809v1 Announce Type: cross Abstract: Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive medi

Today we're releasing ZAYA1-8B, a reasoning MoE trained on @AMD and optimized for intelligence density. With <1B active params, it outperfor…

Model ReleasesDGX agent

Today we're releasing ZAYA1-8B, a reasoning MoE trained on @AMD and optimized for intelligence density. With <1B active params, it outperforms open-weight models many times its size on math and reason

Toward a Science of Intent: Closure Gaps and Delegation Envelopes for Open-World AI Agents

Model ReleasesDGX agent

arXiv:2604.25000v2 Announce Type: replace Abstract: Recent work has framed intelligence in verifiable tasks as reducing time-to-solution through learned structure and test-time search, while systems w

Toward Generative Quantum Utility via Correlation-Complexity Map

Model ReleasesDGX agent

arXiv:2603.06440v2 Announce Type: replace Abstract: We study a practical question in generative quantum machine learning: given a classical dataset, can we determine, before training, whether it is we

Towards Agentic Runtime Healing

Model ReleasesDGX agent

arXiv:2408.01055v2 Announce Type: replace-cross Abstract: Self-healing systems have long been a focus of research, aiming to enable software to recover from unexpected runtime errors without human int

Towards Multi-Agent Autonomous Reasoning in Hydrodynamics

Model ReleasesDGX agent

arXiv:2605.01102v1 Announce Type: new Abstract: Single-agent systems (SAS) have become the default pattern for LLM-driven scientific workflows, but routing planning, tool use, and synthesis through a

Towards Understanding Specification Gaming in Reasoning Models

Model ReleasesDGX agent

arXiv:2605.02269v1 Announce Type: new Abstract: Specification gaming is a critical failure mode of LLM agents. Despite this, there has been little systematic research into when it arises and what driv

Training-Free Probabilistic Time-Series Forecasting with Conformal Seasonal Pools

Model ReleasesDGX agent

arXiv:2605.03789v1 Announce Type: cross Abstract: We propose Conformal Seasonal Pools (CSP), a training-free probabilistic time-series forecaster that mixes same-season empirical draws with signed res

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

Model ReleasesDGX agent

arXiv:2605.03792v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar e

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration

Model ReleasesDGX agent

arXiv:2605.01970v2 Announce Type: cross Abstract: Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characte

Two frontier labs. One accelerated computing platform. Congrats to @SpaceX and @AnthropicAI on the new compute partnership, powered by 220,0…

Model ReleasesDGX agent

Two frontier labs. One accelerated computing platform. Congrats to @SpaceX and @AnthropicAI on the new compute partnership, powered by 220,000+ NVIDIA GPUs inside Colossus 1. The future of AI runs on

Two weeks after release, Hy3 preview is #1 on @OpenRouter's weekly leaderboard with 3.66T tokens processed, up 298% week-over-week. #1 in ov…

Model ReleasesDGX agent

Two weeks after release, Hy3 preview is #1 on @OpenRouter's weekly leaderboard with 3.66T tokens processed, up 298% week-over-week. #1 in overall usage, tool calls, and coding. 15.4% market share acro

Uber uses OpenAI to help people earn smarter and book faster

Model ReleasesDGX agent

Uber has integrated OpenAI's technology to enhance its platform, helping drivers optimize earnings through intelligent features and enabling users to book rides more efficiently. This partnership like

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning

Model ReleasesDGX agent

arXiv:2605.03950v1 Announce Type: new Abstract: Although recent LMMs have become much stronger at visual perception, they remain unreliable on problems that require multi-step reasoning over visual ev

← Previous
1…283284285286287…376
Next →