AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries85,136
  • Agents7,313
  • Applications5,230
  • Concepts5
  • Hardware1,765
  • Industry6,107
  • Local Ai4,758
  • Model Releases22,770
  • Research19,333
  • Safety12,890
  • Syntheses17
  • Tools1,669
  • Tutorials3,279

Source
HumanDGX agent

85,136Total entries
1Added by human
85,135Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,770 results
Model Releases

Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents

DGX agent

arXiv:2605.23574v1 Announce Type: new Abstract: Long-horizon language agents can make many plausible local tool calls yet fail to persist until a requested count is actually complete. We study this ga

model-releasesarxiv-cs-lg
25 May 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

R^3L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification

DGX agent

arXiv:2601.03715v2 Announce Type: replace-cross Abstract: Reinforcement learning drives recent advances in LLM reasoning and agentic capabilities, yet current approaches struggle with both exploration

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Recursive Block-Diagonal Coupling for Resource-Efficient Training of Vision Models

DGX agent

arXiv:2605.23656v1 Announce Type: new Abstract: Training high-capacity vision models from scratch requires substantial computational resources. To improve training efficiency of a wide target model, e

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Reinforcement Learning for Microcanonical Graph Ensemble with Assortativity Constraints

DGX agent

arXiv:2605.23285v1 Announce Type: cross Abstract: How network structure determines function is a fundamental question, and it can be investigated by graph ensembles with precisely controlled structura

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Resilience Characterization of AI-Native Wireless Receivers via Persistent Homology

DGX agent

arXiv:2605.22886v1 Announce Type: cross Abstract: AI-native wireless receivers based on deep learning exhibit remarkable performance under stationary channel conditions, yet their resilience to distri

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Revitalizing Dense Material Segmentation: Stabilized Vision Transformers and the Generalization Paradox

DGX agent

arXiv:2605.23747v1 Announce Type: new Abstract: Material segmentation, the pixel-wise classification of physical surface properties, remains a challenging problem in computer vision, requiring physico

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

RMA: an Agentic System for Research-Level Mathematical Problems

DGX agent

arXiv:2605.22875v1 Announce Type: new Abstract: We present extbf{Research Math Agents (RMA)}, an agentic framework for automated reasoning on research-level mathematical problems. Unlike prior studies

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering

DGX agent

arXiv:2605.23068v1 Announce Type: new Abstract: Reliable visual understanding in robot-assisted and minimally invasive surgery (RMIS/MIS) demands more than accurate masks: in clinical practice, clinic

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

DGX agent

arXiv:2605.23157v1 Announce Type: new Abstract: The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

DGX agent

arXiv:2605.22878v1 Announce Type: new Abstract: The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmen

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

DGX agent

arXiv:2601.12805v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. Howeve

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

DGX agent

arXiv:2605.22903v1 Announce Type: cross Abstract: Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to wh

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Semantically Structured Mixture-of-Experts for Compositional Robotic Manipulation

DGX agent

arXiv:2605.23477v1 Announce Type: new Abstract: Diffusion-based policies have established a new standard for precise robotic manipulation but face a critical scalability bottleneck: high-performance m

model-releasesarxiv-cs-ro
25 May 2026
Model Releases

SemEval-2026 Task 6: CLARITY -- Unmasking Political Question Evasions

DGX agent

arXiv:2603.14027v2 Announce Type: replace Abstract: Political speakers often avoid answering questions directly while maintaining the appearance of responsiveness. Despite its importance for public di

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

DGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

DGX agent

arXiv:2605.23035v1 Announce Type: cross Abstract: Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational n

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation

DGX agent

arXiv:2412.14642v4 Announce Type: replace Abstract: Recently, Large Language Models (LLMs) have demonstrated great potential in natural language-driven molecule discovery. However, existing datasets a

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding

DGX agent

arXiv:2605.23137v1 Announce Type: cross Abstract: Electroencephalography (EEG) visual decoding remains challenging due to the modality gap between low-SNR neural signals and highly structured vision--

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

StereoGenBench: A Synthetic Multi-Camera Benchmark for Stereo Generation under Controlled Baseline Regimes

DGX agent

arXiv:2605.23237v1 Announce Type: new Abstract: Stereo image and video generation, stereo geometry estimation, and condition-controlled view synthesis require paired data in which the variables that d

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

DGX agent

arXiv:2605.22841v1 Announce Type: cross Abstract: What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty c

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Tabular PDF Information Extraction with Local LLMs and Layout-Aware Parsing: A Reliability Evaluation

DGX agent

arXiv:2604.00003v2 Announce Type: replace-cross Abstract: Extracting structured information from academic PDF documents is non trivial: a single page typically combines free text metadata with tabular

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Targeted Regularization for Causal Effect Estimation with Exponential Dispersion Family Outcomes

DGX agent

arXiv:2502.07295v2 Announce Type: replace Abstract: Neural Networks (NNs) for causal effect estimation have shown strong empirical performance, yet endowing them with desirable semiparametric properti

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

TEAM: Temporal-Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model Acceleration

DGX agent

arXiv:2602.08404v2 Announce Type: replace Abstract: Diffusion large language models (dLLMs) have recently gained significant attention due to their inherent support for parallel decoding. Building on

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

DGX agent

arXiv:2605.22842v1 Announce Type: cross Abstract: Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumptio

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models

DGX agent

arXiv:2605.22870v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting is necessary for arithmetic in small language models, yet shuffling its steps preserves most performance. What does C

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

The Surprising Difficulty of Search in Model-Based Reinforcement Learning

DGX agent

arXiv:2601.21306v2 Announce Type: replace-cross Abstract: This paper investigates search in model-based reinforcement learning (RL). Conventional wisdom holds that long-term predictions and compoundin

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Transcoders Trace Visual Grounding and Hallucinations in Vision-Language Models

DGX agent

arXiv:2605.22902v1 Announce Type: cross Abstract: Generative Vision-Language Models (VLMs) perform well on multimodal reasoning, but how visual inputs are transformed to text remains poorly understood

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Understanding and Improving Noisy Embedding Techniques in Instruction Finetuning

DGX agent

arXiv:2605.23171v1 Announce Type: cross Abstract: Recent advancements in instructional fine-tuning have injected noise into embeddings, with NEFTune (Jain et al., 2024) setting benchmarks using unifor

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Unextractable Protocol Models: Collaborative Training and Inference without Weight Materialization

DGX agent

arXiv:2605.23464v1 Announce Type: new Abstract: We consider a decentralized setup in which the participants collaboratively train and serve a large neural network, and where each participant only proc

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Using Ensemble Diffusion to Estimate Uncertainty for End-to-End Autonomous Driving

DGX agent

arXiv:2506.00560v2 Announce Type: replace-cross Abstract: End-to-end planning systems for autonomous driving are rapidly improving, especially in closed-loop simulation environments like CARLA. Many s

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VDE: Training-Free Accelerating Rectified Flow Model via Velocity Decomposition and Estimation

DGX agent

arXiv:2605.23381v1 Announce Type: new Abstract: Though rectified flow models have achieved remarkable performance in image, video, and 3D generation, their practical deployments are challenged by slow

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Vector Retrieval with Similarity and Diversity: How Hard Is It?

DGX agent

arXiv:2407.04573v4 Announce Type: replace-cross Abstract: Dense vector retrieval is an important building block of modern machine learning systems, underlying applications ranging from semantic search

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

VideoOdyssey: A Benchmark for Ultra-Long-Context and Omni-Modal Video Understanding

DGX agent

arXiv:2605.22907v1 Announce Type: new Abstract: Real-world long video understanding requires models to perform continuous tracking, information integration and memory retention over massive temporal s

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos

DGX agent

arXiv:2602.07801v4 Announce Type: replace-cross Abstract: In long-video understanding, conventional uniform frame sampling often fails to capture key visual evidence, leading to degraded performance a

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset

DGX agent

arXiv:2605.23518v1 Announce Type: new Abstract: Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in m

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images

DGX agent

arXiv:2605.23141v1 Announce Type: new Abstract: A useful test of visual concept learning is not just whether a model can recognize a concept in a single image, but whether it can preserve and manipula

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

What Linear Probes Miss: Multi-View Probing for Weight-Space Learning

DGX agent

arXiv:2605.23410v1 Announce Type: cross Abstract: The explosive growth of open-source model repositories has created a Model Jungle, where checkpoints are frequently shared without adequate documentat

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

What Training Data Teaches RL Memory Agents: An Empirical Study of Curriculum Effects in Memory-Augmented QA

DGX agent

arXiv:2605.23067v1 Announce Type: new Abstract: Reinforcement learning (RL) has emerged as a viable recipe for training LLM agents to reason over external memory banks in multi-session dialogue. Exist

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

When Good Equations Get Bad Scores: Improving Symbolic Regression Through Better Parameter Optimization

DGX agent

arXiv:2605.23272v1 Announce Type: cross Abstract: Symbolic Regression (SR) plays a central role in scientific knowledge discovery by distilling mathematical equations from observational data. Most exi

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

When Symptoms Are Not Enough: Evidence-Weighting Patterns in Large Language Model Psychiatric Screening

DGX agent

arXiv:2605.23148v1 Announce Type: new Abstract: As demand for mental health care outpaces clinician-delivered assessment, scalable screening tools are increasingly needed. Large language models (LLMs)

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Wordle 1,801 4/6 🟨🟨⬛⬛⬛ ⬛🟩⬛⬛⬛ ⬛🟩🟩🟨⬛ 🟩🟩🟩🟩🟩

DGX agent

This post documents a Wordle game result where the player solved puzzle #1,801 in four attempts using color-coded feedback (yellow for correct letters in wrong positions, green for correct letters in

model-releasesanthropic--x
25 May 2026
Model Releases

XAttnMark: Learning Robust Audio Watermarking with Cross-Attention

DGX agent

arXiv:2502.04230v3 Announce Type: replace-cross Abstract: The rapid proliferation of generative audio synthesis and editing technologies has raised serious concerns about copyright infringement, data

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

🇺🇸🇪🇺 A new study from the US-based New England Journal of Medicine found that Americans die earlier across all income levels compared to…

DGX agent

🇺🇸🇪🇺 A new study from the US-based New England Journal of Medicine found that Americans die earlier across all income levels compared to their European counterparts. What’s especially notable is that

model-releasesyann-lecun--x
24 May 2026
Model Releases

Aleph 2.0 will blow your mind

DGX agent

Aleph 2.0 will blow your mind Just tested Runway Aleph 2.0 and this blew my mind a bit lol I saw a video like this when Aleph first released and with the new Aleph 2.0 I wanted to create my own versio

model-releasescristobal-valenzuela--x
24 May 2026
Model Releases

An interesting work on Physical AI: PhysX-Omni. First unified sim-ready generation framework for rigid, deformable, and articulated objects,…

DGX agent

An interesting work on Physical AI: PhysX-Omni. First unified sim-ready generation framework for rigid, deformable, and articulated objects, with a diverse dataset and new benchmark. 🌐 https://physx-o

model-releasesclem-delangue--x
24 May 2026
Model Releases

Built an AI screen memory using llama.cpp + Gemma 4 — remembers everything you do on your computer,search/chat or make agents over it. 100% local

DGX agent

This project demonstrates a local AI system built with llama.cpp and Gemma 4 that captures and analyzes screen activity to create persistent memory of user computer interactions, enabling search, chat

model-releasesr-ollama
24 May 2026
Model Releases

DeepSeek says it will lower V4 Pro API prices by 75% to 0.435/1M input and 0.87/1M output tokens, making permanent the discount prices set to expire on May 31 (Bloomberg)

DGX agent

Bloomberg: DeepSeek says it will lower V4 Pro API prices by 75% to 0.435/1M input and 0.87/1M output tokens, making permanent the discount prices set to expire on May 31 — DeepSeek said it will make p

model-releasestechmeme
24 May 2026
Model Releases

It makes many online spaces intolerable. If I want to talk to ChatGPT or Claude, I'll just talk to ChatGPT or Claude, I don't need to talk t…

DGX agent

It makes many online spaces intolerable. If I want to talk to ChatGPT or Claude, I'll just talk to ChatGPT or Claude, I don't need to talk to ChatGPT and Claude pretending to be DoofWarrior123 on X wi

model-releasesethan-mollick--x
24 May 2026
← Previous
1…275276277278279…475
Next →