AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
All
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,545 results
Model Releases

WaferSAGE: Large Language Model-Powered Wafer Defect Analysis via Synthetic Data Generation and Rubric-Guided Reinforcement Learning

DGX agent

arXiv:2604.27629v1 Announce Type: new Abstract: We present WaferSAGE, a framework for wafer defect visual question answering using small vision-language models. To address data scarcity in semiconduct

model-releasesarxiv-cs-ai
1 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

WaPo: Trump’s border wall expansion has bulldozed an ancient tribal site in the Arizona desert—damaging a massive Indigenous ground etching …

DGX agent

WaPo: Trump’s border wall expansion has bulldozed an ancient tribal site in the Arizona desert—damaging a massive Indigenous ground etching that is believed to be at least 1,000 years old. https://www

model-releasesanthropic--x
1 May 2026
Model Releases

WebMall -- A Multi-Shop Benchmark for Evaluating Web Agents

DGX agent

arXiv:2508.13024v3 Announce Type: replace Abstract: LLM-based web agents have the potential to automate long-running web tasks, such as searching for products in multiple e-shops and subsequently orde

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

What benchmark would you build for “reply quality” in SDR generation? [D]

DGX agent

Based on the Reddit discussion title, this likely discusses how to design and establish benchmarks for evaluating the quality of AI-generated responses in Sales Development Representative (SDR) system

model-releasesr-machinelearning
1 May 2026
Model Releases

What Makes a Good Terminal-Agent Benchmark Task: A Guideline for Adversarial, Difficult, and Legible Evaluation Design

DGX agent

arXiv:2604.28093v1 Announce Type: new Abstract: Terminal-agent benchmarks have become a primary signal for measuring the coding and system-administration capabilities of large language models. As the

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

What Suppresses Nash Equilibrium Play in Large Language Models? Mechanistic Evidence and Causal Control

DGX agent

arXiv:2604.27167v1 Announce Type: cross Abstract: LLM agents are known to deviate from Nash equilibria in strategic interactions, but nobody has looked inside the model to understand why, or asked whe

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

When Continual Learning Moves to Memory: A Study of Experience Reuse in LLM Agents

DGX agent

arXiv:2604.27003v1 Announce Type: cross Abstract: Memory-augmented LLM agents offer an appealing shortcut to continual learning: rather than updating model parameters, they accumulate experience in ex

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis

DGX agent

arXiv:2604.27228v1 Announce Type: new Abstract: Democratic discourse analysis systems increasingly rely on multi-agent LLM pipelines in which distinct evaluator models are assigned adversarial roles t

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable,…

DGX agent

When training Grok 4.3, we spoke directly with devs and businesses to understand what they actually needed: a model that’s fast, affordable, and great at tool calling. The result is a daily driver tha

model-releaseselon-musk--x
1 May 2026
Model Releases

WindowsWorld: A Process-Centric Benchmark of Autonomous GUI Agents in Professional Cross-Application Environments

DGX agent

arXiv:2604.27776v1 Announce Type: new Abstract: While GUI agents have shown impressive capabilities in common computer-use tasks such as OSWorld, current benchmarks mainly focus on isolated and single

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Wordle 1,776 4/6 ⬛⬛⬛🟨⬛ ⬛⬛⬛⬛🟨 ⬛🟩🟩⬛🟩 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle game result where the player solved puzzle #1,776 in 4 attempts, using color-coded emoji feedback to show which letters were correct, misplaced, or incorrect with each gue

model-releasesanthropic--x
1 May 2026
Model Releases

Wordle 1,777 5/6 ⬛⬛⬛⬛🟩 🟨⬛⬛⬛⬛ ⬛⬛🟨⬛🟩 ⬛🟩🟩🟩🟩 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle game result (puzzle 1,777) where the player achieved a solution in 5 attempts, using the standard color-coded feedback system (gray for incorrect letters, yellow for corre

model-releasesanthropic--x
1 May 2026
Model Releases

WST-X Series: Wavelet Scattering Transform for Interpretable Speech Deepfake Detection

DGX agent

arXiv:2602.02980v2 Announce Type: replace-cross Abstract: In this work, we focus on front-end design for speech deepfake detectors, the component that determines the discriminative acoustic cues provi

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

xAI has Released Voice Cloning in API Console in US 🇺🇸

DGX agent

xAI has released a voice cloning feature in its API Console, currently available to users in the United States. This capability allows developers to create synthetic voices based on audio samples thro

model-releaseselon-musk--x
1 May 2026
Model Releases

xAI launches Grok 4.3, featuring 'always-on reasoning', 1M token context window, and low API pricing, and releases a voice cloning suite called Custom Voices (Carl Franzen/VentureBeat)

DGX agent

Carl Franzen / VentureBeat: xAI launches Grok 4.3, featuring “always-on reasoning”, 1M token context window, and low API pricing, and releases a voice cloning suite called Custom Voices — While Elon M

model-releasestechmeme
1 May 2026
Model Releases

you know what all of these 'which is better' polls are silly use codex or claude code, whatever works best for you i am grateful we live in …

DGX agent

you know what all of these 'which is better' polls are silly use codex or claude code, whatever works best for you i am grateful we live in a time with such amazing tools, and grateful there is a choi

model-releasessam-altman--x
1 May 2026
Model Releases

3D-LENS: A 3D Lifting-based Elevated Novel-view Synthesis method for Single-View Aerial-Ground Re-Identification

DGX agent

arXiv:2604.26520v1 Announce Type: new Abstract: Aerial-Ground Re-Identification (AG-ReID) is constrained by the viewpoint-domain gap, as drastic viewpoint disparities occlude or distort discriminative

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

5/5 Maestro automatically searches across this space (model ensembles, scaling strategies, execution policies), maps the Pareto frontier, an…

DGX agent

5/5 Maestro automatically searches across this space (model ensembles, scaling strategies, execution policies), maps the Pareto frontier, and surfaces the full accuracy–cost–latency tradeoff surface.

model-releasesai21-labs--x
30 Apr 2026
Model Releases

A Multi-Dataset Benchmark of Multiple Instance Learning for 3D Neuroimage Classification

DGX agent

arXiv:2604.26807v1 Announce Type: new Abstract: Despite being resource-intensive to train, 3D convolutional neural networks (CNNs) have been the standard approach to classify CT and MRI scans. Recent

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

A Multistage Extraction Pipeline for Long Scanned Financial Documents: An Empirical Study in Industrial KYC Workflows

DGX agent

arXiv:2604.26462v1 Announce Type: new Abstract: Structured information extraction from long, multilingual scanned financial documents is a core requirement in industrial KYC and compliance workflows.

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

A Note on How to Remove the lnln T Term from the Squint Bound

DGX agent

arXiv:2604.26926v1 Announce Type: new Abstract: In Orabona and Pal [2016], we introduced the shifted KT potentials, to remove the ln ln T factor in the parameter-free learning with expert bound. In th

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio

DGX agent

arXiv:2409.06624v4 Announce Type: replace-cross Abstract: Large Language Models (LLM) often need to be Continual Pre-Trained (CPT) to obtain unfamiliar language skills or adapt to new domains. The hug

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

A self-evolving agent for explainable diagnosis of DFT-experiment band-gap mismatch

DGX agent

arXiv:2604.26703v1 Announce Type: cross Abstract: Standard density functional theory (DFT) routinely misclassifies the electronic ground state of correlated and structurally complex compounds, predict

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

A Systematic Comparison of Prompting and Multi-Agent Methods for LLM-based Stance Detection

DGX agent

arXiv:2604.26319v1 Announce Type: new Abstract: Stance detection identifies the attitude of a text author toward a given target. Recent studies have explored various LLM-based strategies for this task

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

AdaMem: Adaptive User-Centric Memory for Long-Horizon Dialogue Agents

DGX agent

arXiv:2603.16496v2 Announce Type: replace Abstract: Large language model (LLM) agents increasingly rely on external memory to support long-horizon interaction, personalized assistance, and multi-step

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

Adaptive and Fine-grained Module-wise Expert Pruning for Efficient LoRA-MoE Fine-Tuning

DGX agent

arXiv:2604.26340v1 Announce Type: new Abstract: LoRA-MoE has emerged as an effective paradigm for parameter-efficient fine-tuning, combining the low training cost of LoRA with the increased adaptation

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

Adaptive Scaling of Policy Constraints for Offline Reinforcement Learning

DGX agent

arXiv:2508.19900v2 Announce Type: replace Abstract: Offline reinforcement learning (RL) enables learning effective policies from fixed datasets without any environment interaction. Existing methods ty

model-releasesarxiv-cs-lg
30 Apr 2026
Model Releases

Affective Flow Language Model for Emotional Support Conversation

DGX agent

arXiv:2602.08826v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have been widely applied to emotional support conversation (ESC). However, complex multi-turn support remains cha

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Agentic Search in the Wild: Intents and Trajectory Dynamics from 14M+ Real Search Requests

DGX agent

arXiv:2601.17617v3 Announce Type: replace-cross Abstract: LLM-powered search agents are increasingly being used for multi-step information seeking tasks, yet the IR community lacks empirical understan

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

音声AIの素早さと賢さを両立できるか? 私たち人間は会話の中で、言いたいことを全部まとめてから話し始めるのではなく、話しながら考えを整理していきます。応答の速い Speech-to-Speech モデルは、この「話しながら考える」を実現しましたが、そのぶん思考が浅くなりがちです。…

DGX agent

音声AIの素早さと賢さを両立できるか? 私たち人間は会話の中で、言いたいことを全部まとめてから話し始めるのではなく、話しながら考えを整理していきます。応答の速い Speech-to-Speech モデルは、この「話しながら考える」を実現しましたが、そのぶん思考が浅くなりがちです。かといって知識豊富な LLM を挟むカスケード型では、遅延が生じるため「話しながら」が成立しません。 そこで Sakan

model-releasesdavid-ha--x
30 Apr 2026
Model Releases

AirZoo: A Unified Large-Scale Dataset for Grounding Aerial Geometric 3D Vision

DGX agent

arXiv:2604.26567v1 Announce Type: new Abstract: Despite the rapid progress in data-driven 3D vision, aerial geometric 3D vision remains a formidable challenge due to the severe scarcity of large-scale

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

alignment failure

DGX agent

alignment failure Fun fact - if you have a recent commit that mentions OpenClaw in a json blob, Claude Code will either refuse your request or bill you extra money. This is an empty repo, I'm just cal

model-releasessam-altman--x
30 Apr 2026
Model Releases

Allen AI just released the OlmPool research series on Hugging Face Early 7-8B checkpoints trained to 150B tokens exploring how minor archite…

DGX agent

Allen AI released the OlmPool research series on Hugging Face, featuring early 7-8B parameter language model checkpoints trained on 150 billion tokens. The research explores how minor architectural mo

model-releasesclem-delangue--x
30 Apr 2026
Model Releases

Anthropic announces Claude Security public beta to find and fix software vulnerabilities

DGX agent

Anthropic PBC announced the launch of Claude Security in public beta mode today to help cybersecurity teams scan their codebases for vulnerabilities and generate patches. Part of Claude Enterprise, th

model-releasessiliconangle
30 Apr 2026
Model Releases

Anthropic unveils BioMysteryBench to test Claude's bioinformatics skills against human experts, and says Mythos solved ~30% of 23 questions that stumped experts (Anthropic)

DGX agent

Anthropic: Anthropic unveils BioMysteryBench to test Claude's bioinformatics skills against human experts, and says Mythos solved ~30% of 23 questions that stumped experts — In this post, Brianna, a r

model-releasestechmeme
30 Apr 2026
Model Releases

APEX-Agents now has a @huggingface leaderboard for open-source models. APEX-Agents is our frontier benchmark for whether models can do the r…

DGX agent

APEX-Agents now has a @huggingface leaderboard for open-source models. APEX-Agents is our frontier benchmark for whether models can do the real work of consultants, lawyers, and bankers. https://huggi

model-releasesclem-delangue--x
30 Apr 2026
Model Releases

Associative-State Universal Transformers: Sparse Retrieval Meets Structured Recurrence

DGX agent

arXiv:2604.25930v1 Announce Type: new Abstract: We study whether a structured recurrent state can serve as a compact associative backbone for language modeling while still supporting exact retrieval.

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

Auditing Marketing Budget Allocation with Hindsight Regret

DGX agent

arXiv:2604.25977v1 Announce Type: cross Abstract: Organizations routinely make strategic budget allocations under operational constraints, but often lack a principled way to assess whether realized al

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what ou…

DGX agent

> be me > 'the internet is polluted by ai slop, we need low-background tokens' > 'wouldnt it be cool if we could time travel and see what our ancestors 100 years ago would say to us' > all the existin

model-releasesswyx--x
30 Apr 2026
Model Releases

Benchmarking Complex Multimodal Document Processing Pipelines: A Unified Evaluation Framework for Enterprise AI

DGX agent

arXiv:2604.26382v1 Announce Type: cross Abstract: Most enterprise document AI today is a pipeline. Parse, index, retrieve, generate. Each of those stages has been studied to death on its own -- what's

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Benchmarking Deep Learning and Vision Foundation Models for Atypical vs. Normal Mitosis Classification with Cross-Dataset Evaluation

DGX agent

arXiv:2506.21444v4 Announce Type: replace Abstract: Atypical mitosis marks a deviation in the cell division process that has been shown be an independent prognostic marker for tumor malignancy. Howeve

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-Codex

DGX agent

arXiv:2604.14858v2 Announce Type: replace Abstract: As agent systems move into increasingly diverse execution settings, trajectory-level safety evaluation and diagnosis require benchmarks that evolve

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Beyond Fixed Formulas: Data-Driven Linear Predictor for Efficient Diffusion Models

DGX agent

arXiv:2604.26365v1 Announce Type: new Abstract: To address the high sampling cost of Diffusion Transformers (DiTs), feature caching offers a training-free acceleration method. However, existing method

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

Beyond the Leaderboard: Rethinking Medical Benchmarks for Large Language Models

DGX agent

arXiv:2508.04325v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show significant potential in healthcare, prompting numerous benchmarks to evaluate their capabilities. However,

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Breaking the Rigid Prior: Towards Articulated 3D Anomaly Detection

DGX agent

arXiv:2604.26868v1 Announce Type: new Abstract: Existing 3D anomaly detection methods are built on a rigid prior: normal geometry is pose-invariant and can be canonicalized through registration or ali

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

Bridge: Basis-Driven Causal Inference Marries VFMs for Domain Generalization

DGX agent

arXiv:2604.26820v1 Announce Type: new Abstract: Detectors often suffer from degraded performance, primarily due to the distributional gap between the source and target domains. This issue is especiall

model-releasesarxiv-cs-cv
30 Apr 2026
Model Releases

Calibrated Surprise: An Information-Theoretic Account of Creative Quality

DGX agent

arXiv:2604.26269v1 Announce Type: cross Abstract: The essence of good creative writing is calibrated surprise: when constraints from all relevant dimensions act together, the feasible solution space c

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems

DGX agent

arXiv:2604.26145v1 Announce Type: cross Abstract: AI-powered language learning tools increasingly provide instant, personalised feedback to millions of learners worldwide. However, this feedback can f

model-releasesarxiv-cs-ai
30 Apr 2026
← Previous
1…370371372373374…470
Next →