AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
All
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,620 results
Model Releases

RMA: an Agentic System for Research-Level Mathematical Problems

DGX agent

arXiv:2605.22875v1 Announce Type: new Abstract: We present extbf{Research Math Agents (RMA)}, an agentic framework for automated reasoning on research-level mathematical problems. Unlike prior studies

model-releasesarxiv-cs-ai
25 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering

DGX agent

arXiv:2605.23068v1 Announce Type: new Abstract: Reliable visual understanding in robot-assisted and minimally invasive surgery (RMIS/MIS) demands more than accurate masks: in clinical practice, clinic

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

DGX agent

arXiv:2605.23157v1 Announce Type: new Abstract: The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

DGX agent

arXiv:2605.22878v1 Announce Type: new Abstract: The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmen

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

DGX agent

arXiv:2601.12805v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. Howeve

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

DGX agent

arXiv:2605.22903v1 Announce Type: cross Abstract: Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to wh

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Semantically Structured Mixture-of-Experts for Compositional Robotic Manipulation

DGX agent

arXiv:2605.23477v1 Announce Type: new Abstract: Diffusion-based policies have established a new standard for precise robotic manipulation but face a critical scalability bottleneck: high-performance m

model-releasesarxiv-cs-ro
25 May 2026
Model Releases

SemEval-2026 Task 6: CLARITY -- Unmasking Political Question Evasions

DGX agent

arXiv:2603.14027v2 Announce Type: replace Abstract: Political speakers often avoid answering questions directly while maintaining the appearance of responsiveness. Despite its importance for public di

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

DGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

DGX agent

arXiv:2605.23035v1 Announce Type: cross Abstract: Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational n

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation

DGX agent

arXiv:2412.14642v4 Announce Type: replace Abstract: Recently, Large Language Models (LLMs) have demonstrated great potential in natural language-driven molecule discovery. However, existing datasets a

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding

DGX agent

arXiv:2605.23137v1 Announce Type: cross Abstract: Electroencephalography (EEG) visual decoding remains challenging due to the modality gap between low-SNR neural signals and highly structured vision--

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

StereoGenBench: A Synthetic Multi-Camera Benchmark for Stereo Generation under Controlled Baseline Regimes

DGX agent

arXiv:2605.23237v1 Announce Type: new Abstract: Stereo image and video generation, stereo geometry estimation, and condition-controlled view synthesis require paired data in which the variables that d

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

DGX agent

arXiv:2605.22841v1 Announce Type: cross Abstract: What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty c

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation

DGX agent

arXiv:2604.00003v2 Announce Type: replace-cross Abstract: Extracting structured information from academic PDF documents is non trivial: a single page typically combines free text metadata with tabular

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Targeted Regularization for Causal Effect Estimation with Exponential Dispersion Family Outcomes

DGX agent

arXiv:2502.07295v2 Announce Type: replace Abstract: Neural Networks (NNs) for causal effect estimation have shown strong empirical performance, yet endowing them with desirable semiparametric properti

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration

DGX agent

arXiv:2602.08404v2 Announce Type: replace Abstract: Diffusion large language models (dLLMs) have recently gained significant attention due to their inherent support for parallel decoding. Building on

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

DGX agent

arXiv:2605.22842v1 Announce Type: cross Abstract: Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumptio

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models

DGX agent

arXiv:2605.22870v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting is necessary for arithmetic in small language models, yet shuffling its steps preserves most performance. What does C

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

The Surprising Difficulty of Search in Model-Based Reinforcement Learning

DGX agent

arXiv:2601.21306v2 Announce Type: replace-cross Abstract: This paper investigates search in model-based reinforcement learning (RL). Conventional wisdom holds that long-term predictions and compoundin

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

DGX agent

arXiv:2605.22902v1 Announce Type: cross Abstract: Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning

DGX agent

arXiv:2605.23171v1 Announce Type: cross Abstract: Recent advancements in instructional fine-tuning have injected noise into embeddings, with NEFTune (Jain et al., 2024) setting benchmarks using unifor

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization

DGX agent

arXiv:2605.23464v1 Announce Type: new Abstract: We consider a decentralized setup in which the participants collaboratively train and serve a large neural network, and where each participant only proc

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Using Ensemble Diffusion to Estimate Uncertainty for End-to-End Autonomous Driving

DGX agent

arXiv:2506.00560v2 Announce Type: replace-cross Abstract: End-to-end planning systems for autonomous driving are rapidly improving, especially in closed-loop simulation environments like CARLA. Many s

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation

DGX agent

arXiv:2605.23381v1 Announce Type: new Abstract: Though rectified flow models have achieved remarkable performance in image, video, and 3D generation, their practical deployments are challenged by slow

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Vector Retrieval with Similarity and Diversity: How Hard Is It?

DGX agent

arXiv:2407.04573v4 Announce Type: replace-cross Abstract: Dense vector retrieval is an important building block of modern machine learning systems, underlying applications ranging from semantic search

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

DGX agent

arXiv:2605.22907v1 Announce Type: new Abstract: Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal s

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

DGX agent

arXiv:2602.07801v4 Announce Type: replace-cross Abstract: In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance a

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

DGX agent

arXiv:2605.23518v1 Announce Type: new Abstract: Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in m

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

DGX agent

arXiv:2605.23141v1 Announce Type: new Abstract: A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipula

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

What Linear Probes Miss: Multi-View Probing for Weight-Space Learning

DGX agent

arXiv:2605.23410v1 Announce Type: cross Abstract: The explosive growth of open-source model repositories has created a Model Jungle, where checkpoints are frequently shared without adequate documentat

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA

DGX agent

arXiv:2605.23067v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a viable recipe for training LLM agents to reason over external memory banks in multi-session dialogue. Exist

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization

DGX agent

arXiv:2605.23272v1 Announce Type: cross Abstract: Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational data. Most exi

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening

DGX agent

arXiv:2605.23148v1 Announce Type: new Abstract: As demand for mental health care outpaces clinician-delivered assessment, scalable screening tools are increasingly needed. Large language models (LLMs)

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Wordle 1,801 4/6 🟨🟨⬛⬛⬛ ⬛🟩⬛⬛⬛ ⬛🟩🟩🟨⬛ 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle game result where the player solved puzzle #1,801 in four attempts using color-coded feedback (yellow for correct letters in wrong positions, green for correct letters in

model-releasesanthropic--x
25 May 2026
Model Releases

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention

DGX agent

arXiv:2502.04230v3 Announce Type: replace-cross Abstract: The rapid proliferation of generative audio synthesis and editing technologies has raised serious concerns about copyright infringement, data

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

🇺🇸🇪🇺 A new study from the US-based New England Journal of Medicine found that Americans die earlier across all income levels compared to…

DGX agent

🇺🇸🇪🇺 A new study from the US-based New England Journal of Medicine found that Americans die earlier across all income levels compared to their European counterparts. What’s especially notable is that

model-releasesyann-lecun--x
24 May 2026
Model Releases

Aleph 2.0 will blow your mind

DGX agent

Aleph 2.0 will blow your mind Just tested Runway Aleph 2.0 and this blew my mind a bit lol I saw a video like this when Aleph first released and with the new Aleph 2.0 I wanted to create my own versio

model-releasescristobal-valenzuela--x
24 May 2026
Model Releases

An interesting work on Physical AI: PhysX-Omni. First unified sim-ready generation framework for rigid, deformable, and articulated objects,…

DGX agent

An interesting work on Physical AI: PhysX-Omni. First unified sim-ready generation framework for rigid, deformable, and articulated objects, with a diverse dataset and new benchmark. 🌐 https://physx-o

model-releasesclem-delangue--x
24 May 2026
Model Releases

Built an AI screen memory using llama.cpp + Gemma 4 — remembers everything you do on your computer,search/chat or make agents over it. 100% local

DGX agent

This project demonstrates a local AI system built with llama.cpp and Gemma 4 that captures and analyzes screen activity to create persistent memory of user computer interactions, enabling search, chat

model-releasesr-ollama
24 May 2026
Model Releases

DeepSeek says it will lower V4 Pro API prices by 75% to 0.435/1M input and 0.87/1M output tokens, making permanent the discount prices set to expire on May 31 (Bloomberg)

DGX agent

Bloomberg: DeepSeek says it will lower V4 Pro API prices by 75% to 0.435/1M input and 0.87/1M output tokens, making permanent the discount prices set to expire on May 31 — DeepSeek said it will make p

model-releasestechmeme
24 May 2026
Model Releases

It makes many online spaces intolerable. If I want to talk to ChatGPT or Claude, I'll just talk to ChatGPT or Claude, I don't need to talk t…

DGX agent

It makes many online spaces intolerable. If I want to talk to ChatGPT or Claude, I'll just talk to ChatGPT or Claude, I don't need to talk to ChatGPT and Claude pretending to be DoofWarrior123 on X wi

model-releasesethan-mollick--x
24 May 2026
Model Releases

It’s no longer just AI companies & their founders being sued over AI training - individual researchers are now being sued, too. In a new law…

DGX agent

It’s no longer just AI companies & their founders being sued over AI training - individual researchers are now being sued, too. In a new lawsuit, two authors allege that Guillaume Lample, while an AI

model-releasesgary-marcus--x
24 May 2026
Model Releases

Mad House — Usborne Creepy Computer Games

DGX agent

Tool: Mad House — Usborne Creepy Computer Games Via Hacker News I learned that UK publisher Usborne published free PDFs of their 1980s Computer Books, some of which I remember working through on my Co

model-releasessimon-willison
24 May 2026
Model Releases

People often ask what my biggest tip is for getting the most out of Claude Code. These days my #1 tip is: use auto mode Auto mode means no m…

DGX agent

People often ask what my biggest tip is for getting the most out of Claude Code. These days my #1 tip is: use auto mode Auto mode means no more permission prompts. It is the key building block for mul

model-releasesboris-cherny--x
24 May 2026
Model Releases

quick summary of someone's github, cool! here's me

DGX agent

quick summary of someone's github, cool! here's me I always wanted a GitHub dashboard: See my repos, open Issues/PRs, what version I released last, how many commits since last release. So I built one

model-releasesyohei-nakajima--x
24 May 2026
Model Releases

Wordle 1,799 4/6 ⬛⬛⬛⬛⬛ ⬛⬛⬛⬛⬛ ⬛🟨⬛⬛⬛ 🟩🟩🟩🟩🟩

DGX agent

This post shows a Wordle game result where the player solved puzzle #1,799 in 4 attempts, with the final answer being a five-letter word where all letters are in the correct positions (indicated by th

model-releasesanthropic--x
24 May 2026
Model Releases

Wordle 1,800 4/6 ⬛⬛⬛⬛🟩 ⬛🟩🟨⬛⬛ 🟩🟩🟨⬛🟩 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle game result where the player solved puzzle #1,800 in 4 attempts, using the color-coded feedback system (gray for wrong letters, yellow for correct letters in wrong positio

model-releasesanthropic--x
24 May 2026
← Previous
1…272273274275276…472
Next →