AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,512 results
1 May 2026

The latest crop of models remains below 1% on ARC-AGI-3 -- for now. Where will the scores be by the end of the year?

Model ReleasesDGX agent

The latest crop of models remains below 1% on ARC-AGI-3 -- for now. Where will the scores be by the end of the year? GPT-5.5 & Opus 4.7 on ARC-AGI-3 - GPT-5.5: 0.43% - Opus 4.7: 0.18% We found 3 failu

The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please…

Model ReleasesDGX agent

The new Grok comes in below the latest Chinese open weights models, Grok 4 was at the frontier when released. (& Artificial Analysis: please stop using GDPval-AA which is not a useful test of anything

The TEA Nets framework combines AI and cognitive network science to model targets, events and actors in text


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2604.27673v1 Announce Type: new Abstract: We introduce Target-Event-Agent Networks (TEA Nets) as a computational framework to extract subjects (``Agents'), verbs (``Events'), and objects (``Targ

Theory Under Construction: Orchestrating Language Models for Research Software Where the Specification Evolves

Model ReleasesDGX agent

arXiv:2604.27209v1 Announce Type: cross Abstract: Large language models can now generate substantial code and draft research text, but research-software projects require more than either artifact alon

To Build or Not to Build? Factors that Lead to Non-Development or Abandonment of AI Systems

Model ReleasesDGX agent

arXiv:2604.28053v1 Announce Type: cross Abstract: Responsible AI research typically focuses on examining the use and impacts of deployed AI systems. Yet, there is currently limited visibility into the

TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering

Model ReleasesDGX agent

arXiv:2604.28076v1 Announce Type: cross Abstract: Large Language Models (LLMs) have advanced Table Question Answering, where most queries can be answered by extracting information or simple aggregatio

Towards All-Day Perception for Off-Road Driving: A Large-Scale Multispectral Dataset and Comprehensive Benchmark

Model ReleasesDGX agent

arXiv:2604.27499v1 Announce Type: new Abstract: Off-road nighttime autonomous driving suffers from unreliable visible-light perception, making infrared modality crucial for accurate freespace detectio

Towards single-shot coherent imaging via overlap-free ptychography

Model ReleasesDGX agent

arXiv:2602.21361v3 Announce Type: replace-cross Abstract: Ptychographic imaging at synchrotron and XFEL sources requires dense overlapping scans, limiting throughput and increasing dose. Extending coh

TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions

Model ReleasesDGX agent

arXiv:2604.27975v1 Announce Type: cross Abstract: Traditional Shot Boundary Detection (SBD) inherently struggles with complex transitions by formulating the task around isolated cut points, frequently

TripVVT: A Large-Scale Triplet Dataset and a Coarse-Mask Baseline for In-the-Wild Video Virtual Try-On

Model ReleasesDGX agent

arXiv:2604.27958v1 Announce Type: new Abstract: Due to the scarcity of large-scale in-the-wild triplet data and the improper use of masks, the performance of video virtual try-on models remains limite

Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

Model ReleasesDGX agent

arXiv:2604.27093v1 Announce Type: cross Abstract: Current LLM safety alignment techniques improve model robustness against adversarial attacks, but overlook whether and how LLMs can recover helpfulnes

VeraRetouch: A Lightweight Fully Differentiable Framework for Multi-Task Reasoning Photo Retouching

Model ReleasesDGX agent

arXiv:2604.27375v1 Announce Type: new Abstract: Reasoning photo retouching has gained significant traction, requiring models to analyze image defects, give reasoning processes, and execute precise ret

VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking

Model ReleasesDGX agent

arXiv:2601.08611v2 Announce Type: replace-cross Abstract: The growing scale of online misinformation urgently demands Automated Fact-Checking (AFC). Existing benchmarks for evaluating AFC systems, how

Visual Analysis of Multi-outcome Causal Graphs

Model ReleasesDGX agent

arXiv:2408.02679v3 Announce Type: replace Abstract: We introduce a visual analysis method for multiple causal graphs with different outcome variables, namely, multi-outcome causal graphs. Multi-outcom

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Model ReleasesDGX agent

arXiv:2604.28185v1 Announce Type: new Abstract: Recent visual generation models have made major progress in photorealism, typography, instruction following, and interactive editing, yet they still str

WaferSAGE: Large Language Model-Powered Wafer Defect Analysis via Synthetic Data Generation and Rubric-Guided Reinforcement Learning

Model ReleasesDGX agent

arXiv:2604.27629v1 Announce Type: new Abstract: We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconduct

WaPo: Trump’s border wall expansion has bulldozed an ancient tribal site in the Arizona desert—damaging a massive Indigenous ground etching …

Model ReleasesDGX agent

WaPo: Trump’s border wall expansion has bulldozed an ancient tribal site in the Arizona desert—damaging a massive Indigenous ground etching that is believed to be at least 1,000 years old. https://www

WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents

Model ReleasesDGX agent

arXiv:2508.13024v3 Announce Type: replace Abstract: LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently orde

What benchmark would you build for “reply quality” in SDR generation? [D]

Model ReleasesDGX agent

Based on the Reddit discussion title, this likely discusses how to design and establish benchmarks for evaluating the quality of AI-generated responses in Sales Development Representative (SDR) system

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

Model ReleasesDGX agent

arXiv:2604.28093v1 Announce Type: new Abstract: Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the

What Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal Control

Model ReleasesDGX agent

arXiv:2604.27167v1 Announce Type: cross Abstract: LLM agents are known to deviate from Nash equilibria in strategic interactions, but nobody has looked inside the model to understand why, or asked whe

When Continual Learning Moves to Memory: A Study of Experience Reuse in LLM Agents

Model ReleasesDGX agent

arXiv:2604.27003v1 Announce Type: cross Abstract: Memory-augmented LLM agents offer an appealing shortcut to continual learning: rather than updating model parameters, they accumulate experience in ex

When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis

Model ReleasesDGX agent

arXiv:2604.27228v1 Announce Type: new Abstract: Democratic discourse analysis systems increasingly rely on multi-agent LLM pipelines in which distinct evaluator models are assigned adversarial roles t

When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable,…

Model ReleasesDGX agent

When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable, and great at tool calling. The result is a daily driver tha

WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments

Model ReleasesDGX agent

arXiv:2604.27776v1 Announce Type: new Abstract: While GUI agents have shown impressive capabilities in common computer-use tasks such as OSWorld, current benchmarks mainly focus on isolated and single

Wordle 1,776 4/6 ⬛⬛⬛🟨⬛ ⬛⬛⬛⬛🟨 ⬛🟩🟩⬛🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result where the player solved puzzle #1,776 in 4 attempts, using color-coded emoji feedback to show which letters were correct, misplaced, or incorrect with each gue

Wordle 1,777 5/6 ⬛⬛⬛⬛🟩 🟨⬛⬛⬛⬛ ⬛⬛🟨⬛🟩 ⬛🟩🟩🟩🟩 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post documents a Wordle game result (puzzle 1,777) where the player achieved a solution in 5 attempts, using the standard color-coded feedback system (gray for incorrect letters, yellow for corre

WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection

Model ReleasesDGX agent

arXiv:2602.02980v2 Announce Type: replace-cross Abstract: In this work, we focus on front-end design for speech deepfake detectors, the component that determines the discriminative acoustic cues provi

xAI has Released Voice Cloning in API Console in US 🇺🇸

Model ReleasesDGX agent

xAI has released a voice cloning feature in its API Console, currently available to users in the United States. This capability allows developers to create synthetic voices based on audio samples thro

xAI launches Grok 4.3, featuring 'always-on reasoning', 1M token context window, and low API pricing, and releases a voice cloning suite called Custom Voices (Carl Franzen/VentureBeat)

Model ReleasesDGX agent

Carl Franzen / VentureBeat: xAI launches Grok 4.3, featuring “always-on reasoning”, 1M token context window, and low API pricing, and releases a voice cloning suite called Custom Voices — While Elon M

you know what all of these 'which is better' polls are silly use codex or claude code, whatever works best for you i am grateful we live in …

Model ReleasesDGX agent

you know what all of these 'which is better' polls are silly use codex or claude code, whatever works best for you i am grateful we live in a time with such amazing tools, and grateful there is a choi

30 Apr 2026

3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification

Model ReleasesDGX agent

arXiv:2604.26520v1 Announce Type: new Abstract: Aerial-Ground Re-Identification (AG-ReID) is constrained by the viewpoint-domain gap, as drastic viewpoint disparities occlude or distort discriminative

5/5 Maestro automatically searches across this space (model ensembles, scaling strategies, execution policies), maps the Pareto frontier, an…

Model ReleasesDGX agent

5/5 Maestro automatically searches across this space (model ensembles, scaling strategies, execution policies), maps the Pareto frontier, and surfaces the full accuracy–cost–latency tradeoff surface.

A Multi-Dataset Benchmark of Multiple Instance Learning for 3D Neuroimage Classification

Model ReleasesDGX agent

arXiv:2604.26807v1 Announce Type: new Abstract: Despite being resource-intensive to train, 3D convolutional neural networks (CNNs) have been the standard approach to classify CT and MRI scans. Recent

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows

Model ReleasesDGX agent

arXiv:2604.26462v1 Announce Type: new Abstract: Structured information extraction from long, multilingual scanned financial documents is a core requirement in industrial KYC and compliance workflows.

A Note on How to Remove the lnln T Term from the Squint Bound

Model ReleasesDGX agent

arXiv:2604.26926v1 Announce Type: new Abstract: In Orabona and Pal [2016], we introduced the shifted KT potentials, to remove the ln ln T factor in the parameter-free learning with expert bound. In th

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio

Model ReleasesDGX agent

arXiv:2409.06624v4 Announce Type: replace-cross Abstract: Large Language Models (LLM) often need to be Continual Pre-Trained (CPT) to obtain unfamiliar language skills or adapt to new domains. The hug

A self-evolving agent for explainable diagnosis of DFT-experiment band-gap mismatch

Model ReleasesDGX agent

arXiv:2604.26703v1 Announce Type: cross Abstract: Standard density functional theory (DFT) routinely misclassifies the electronic ground state of correlated and structurally complex compounds, predict

A Systematic Comparison of Prompting and Multi-Agent Methods for LLM-based Stance Detection

Model ReleasesDGX agent

arXiv:2604.26319v1 Announce Type: new Abstract: Stance detection identifies the attitude of a text author toward a given target. Recent studies have explored various LLM-based strategies for this task

AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents

Model ReleasesDGX agent

arXiv:2603.16496v2 Announce Type: replace Abstract: Large language model (LLM) agents increasingly rely on external memory to support long-horizon interaction, personalized assistance, and multi-step

Adaptive and Fine-grained Module-wise Expert Pruning for Efficient LoRA-MoE Fine-Tuning

Model ReleasesDGX agent

arXiv:2604.26340v1 Announce Type: new Abstract: LoRA-MoE has emerged as an effective paradigm for parameter-efficient fine-tuning, combining the low training cost of LoRA with the increased adaptation

Adaptive Scaling of Policy Constraints for Offline Reinforcement Learning

Model ReleasesDGX agent

arXiv:2508.19900v2 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables learning effective policies from fixed datasets without any environment interaction. Existing methods ty

Affective Flow Language Model for Emotional Support Conversation

Model ReleasesDGX agent

arXiv:2602.08826v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been widely applied to emotional support conversation (ESC). However, complex multi-turn support remains cha

Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests

Model ReleasesDGX agent

arXiv:2601.17617v3 Announce Type: replace-cross Abstract: LLM-powered search agents are increasingly being used for multi-step information seeking tasks, yet the IR community lacks empirical understan

音声AIの素早さと賢さを両立できるか? 私たち人間は会話の中で、言いたいことを全部まとめてから話し始めるのではなく、話しながら考えを整理していきます。応答の速い Speech-to-Speech モデルは、この「話しながら考える」を実現しましたが、そのぶん思考が浅くなりがちです。…

Model ReleasesDGX agent

音声AIの素早さと賢さを両立できるか? 私たち人間は会話の中で、言いたいことを全部まとめてから話し始めるのではなく、話しながら考えを整理していきます。応答の速い Speech-to-Speech モデルは、この「話しながら考える」を実現しましたが、そのぶん思考が浅くなりがちです。かといって知識豊富な LLM を挟むカスケード型では、遅延が生じるため「話しながら」が成立しません。 そこで Sakan

AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision

Model ReleasesDGX agent

arXiv:2604.26567v1 Announce Type: new Abstract: Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-scale

alignment failure

Model ReleasesDGX agent

alignment failure Fun fact - if you have a recent commit that mentions OpenClaw in a json blob, Claude Code will either refuse your request or bill you extra money. This is an empty repo, I'm just cal

Allen AI just released the OlmPool research series on Hugging Face Early 7-8B checkpoints trained to 150B tokens exploring how minor archite…

Model ReleasesDGX agent

Allen AI released the OlmPool research series on Hugging Face, featuring early 7-8B parameter language model checkpoints trained on 150 billion tokens. The research explores how minor architectural mo

Anthropic announces Claude Security public beta to find and fix software vulnerabilities

Model ReleasesDGX agent

Anthropic PBC announced the launch of Claude Security in public beta mode today to help cybersecurity teams scan their codebases for vulnerabilities and generate patches. Part of Claude Enterprise, th

Anthropic unveils BioMysteryBench to test Claude's bioinformatics skills against human experts, and says Mythos solved ~30% of 23 questions that stumped experts (Anthropic)

Model ReleasesDGX agent

Anthropic: Anthropic unveils BioMysteryBench to test Claude's bioinformatics skills against human experts, and says Mythos solved ~30% of 23 questions that stumped experts — In this post, Brianna, a r

APEX-Agents now has a @huggingface leaderboard for open-source models. APEX-Agents is our frontier benchmark for whether models can do the r…

Model ReleasesDGX agent

APEX-Agents now has a @huggingface leaderboard for open-source models. APEX-Agents is our frontier benchmark for whether models can do the real work of consultants, lawyers, and bankers. https://huggi

Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence

Model ReleasesDGX agent

arXiv:2604.25930v1 Announce Type: new Abstract: We study whether a structured recurrent state can serve as a compact associative backbone for language modeling while still supporting exact retrieval.

Auditing Marketing Budget Allocation with Hindsight Regret

Model ReleasesDGX agent

arXiv:2604.25977v1 Announce Type: cross Abstract: Organizations routinely make strategic budget allocations under operational constraints, but often lack a principled way to assess whether realized al

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what ou…

Model ReleasesDGX agent

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what our ancestors 100 years ago would say to us' > all the existin

Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI

Model ReleasesDGX agent

arXiv:2604.26382v1 Announce Type: cross Abstract: Most enterprise document AI today is a pipeline. Parse, index, retrieve, generate. Each of those stages has been studied to death on its own -- what's

Benchmarking Deep Learning and Vision Foundation Models for Atypical vs. Normal Mitosis Classification with Cross-Dataset Evaluation

Model ReleasesDGX agent

arXiv:2506.21444v4 Announce Type: replace Abstract: Atypical mitosis marks a deviation in the cell division process that has been shown be an independent prognostic marker for tumor malignancy. Howeve

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

Model ReleasesDGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models

Model ReleasesDGX agent

arXiv:2604.26365v1 Announce Type: new Abstract: To address the high sampling cost of Diffusion Transformers (DiTs), feature caching offers a training-free acceleration method. However, existing method

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

Model ReleasesDGX agent

arXiv:2508.04325v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However,

Breaking the Rigid Prior: Towards Articulated 3D Anomaly Detection

Model ReleasesDGX agent

arXiv:2604.26868v1 Announce Type: new Abstract: Existing 3D anomaly detection methods are built on a rigid prior: normal geometry is pose-invariant and can be canonicalized through registration or ali

← Previous
1…295296297298299…376
Next →