AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
Model Releases

BLINKG: A Benchmark for LLM-Integrated Knowledge Graph Generation

DGX agent

arXiv:2605.19518v1 Announce Type: new Abstract: Generating Knowledge Graphs (KGs) remains one of the most time-consuming and labor-intensive tasks for knowledge engineers, as they need to identify sem

model-releasesarxiv-cs-ai
20 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

brain dump of how/why we use Evals to measure agents before & after shipping to prod 1. Good Evals simulate what our real users will do and …

DGX agent

brain dump of how/why we use Evals to measure agents before & after shipping to prod 1. Good Evals simulate what our real users will do and encounter. They’re not really random benchmark tasks, they r

model-releasesharrison-chase--x
20 May 2026
Model Releases

BuildArena: A Physics-Aligned Interactive Benchmark of LLMs for Engineering Construction

DGX agent

arXiv:2510.16559v5 Announce Type: replace Abstract: Engineering construction automation aims to transform natural language specifications into physically viable structures, requiring complex integrate

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Can Large Language Models Reliably Correct Errors in Low-Resource ASR? A Contamination-Aware Case Study on West Frisian

DGX agent

arXiv:2605.19711v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has improved substantially in recent years, yet performance remains limited for low-resource languages. Large languag

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

Can LLMs Emulate Human Belief Dynamics?

DGX agent

arXiv:2605.18781v1 Announce Type: cross Abstract: Can LLMs simulate how humans form and change beliefs in social networks? We put this to the test by replicating an established study on belief dynamic

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

CaptchaMind: Training CAPTCHA Solvers via Reinforcement Learning with Explicit Reasoning Supervision

DGX agent

arXiv:2605.19538v1 Announce Type: cross Abstract: CAPTCHAs are widely deployed as human verification mechanisms and frequently block intelligent agents from completing end-to-end automation in real-wo

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Causal Evidence for Attention Head Imbalance in Modality Conflict Hallucination

DGX agent

arXiv:2605.19250v1 Announce Type: new Abstract: Modality-conflict hallucination occurs when multimodal large language models (MLLMs) prioritize erroneous textual premises over contradictory visual evi

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Chunking German Legal Code

DGX agent

arXiv:2605.19806v1 Announce Type: cross Abstract: This paper investigates chunking strategies for retrieval-augmented generation on German statutory law, using the German Civil Code as a structured be

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning

DGX agent

arXiv:2605.20176v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems have shown promise for clinical decision support, but existing works largely assume that evidence has a

model-releasesarxiv-cs-cl
20 May 2026
Model Releases

ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation

DGX agent

arXiv:2605.18769v1 Announce Type: cross Abstract: Personalized Retrieval-Augmented Generation (RAG) relies on accurately selecting user-relevant documents. In practice, existing RAG approaches often s

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

CogScale: Scalable Benchmark for Sequence Processing

DGX agent

arXiv:2605.19758v1 Announce Type: new Abstract: The ability to maintain and manipulate information over time is a fundamental aspect of living beings and Artificial Intelligence. While modern models h

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

COMPASS: Confined-space Manipulation Planning with Active Sensing Strategy

DGX agent

arXiv:2509.14787v2 Announce Type: replace Abstract: Manipulation in confined and cluttered environments remains a significant challenge due to partial observability and complex configuration spaces. E

model-releasesarxiv-cs-ro
20 May 2026
Model Releases

Compositional Literary Primitives in Instruction-Tuned LLMs: Cross-Architectural SAE Features for Self, Style, and Affect

DGX agent

arXiv:2605.18808v1 Announce Type: cross Abstract: We characterize a compositional architecture of literary primitives in two instruction-tuned large language models (Llama 3.1 8B-Instruct and Gemma 2

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Conflict-Resilient Multi-Agent Reasoning via Signed Graph Modeling

DGX agent

arXiv:2605.19418v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) have demonstrated strong reasoning and decision-making capabilities that consistently surpass those of single LLM ag

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Contrastive Reasoning Alignment: Reinforcement Learning from Hidden Representations

DGX agent

arXiv:2603.17305v2 Announce Type: replace Abstract: We propose CRAFT, a red-teaming alignment framework that leverages model reasoning capabilities and hidden representations to improve robustness aga

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

CRAFT: Critic-Refined Adaptive Key-Frame Targeting for Multimodal Video Question Answering

DGX agent

arXiv:2605.19075v1 Announce Type: cross Abstract: Grounded multi-video question answering over real-world news events requires systems to surface query-relevant evidence across heterogeneous video arc

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Cross-View Attention Fusion Net: A Prior-Guided Dual-View Representation Learning for Cardiac Output Estimation from Short-Term PPG Signals

DGX agent

arXiv:2605.19666v1 Announce Type: cross Abstract: Accurate cardiac output (CO) estimation from photoplethysmography (PPG) is promising for unobtrusive hemodynamic monitoring, but remains difficult sin

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

Cross-View Splatter: Feed-Forward View Synthesis with Georeferenced Images

DGX agent

arXiv:2605.19656v1 Announce Type: new Abstract: We present Cross-View Splatter, a feed-forward method that predicts pixel-aligned Gaussian splats for outdoor scenes captured at ground level AND by sat

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

CutVerse: A Compositional GUI Agents Benchmark for Media Post-Production Editing

DGX agent

arXiv:2605.19484v1 Announce Type: cross Abstract: While GUI agents have made significant progress in web navigation and basic operating system tasks, their capabilities in professional creative workfl

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

D^3-Subsidy: Online and Sequential Driver Subsidy Decision-Making for Large-Scale Ride-Hailing Market

DGX agent

arXiv:2605.20036v1 Announce Type: new Abstract: Ride-hailing platforms like DiDi Chuxing operate in highly dynamic environments where balancing driver supply and passenger demand is critical. Although

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

deadtrees.earth-aerial: A Multi-Resolution Aerial Image Dataset for Tree Cover and Mortality Detection

DGX agent

arXiv:2605.19605v1 Announce Type: new Abstract: Forests worldwide are increasingly threatened by climate change and disturbances such as fire, pests, and pathogens, creating an urgent need for scalabl

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows

DGX agent

arXiv:2605.19099v1 Announce Type: new Abstract: We introduce DecisionBench, a benchmark substrate for emergent delegation in long-horizon agentic workflows. The substrate fixes a task suite (GAIA, tau

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Depth2Pose: A Pose-Based Benchmark for Monocular Depth Estimation without Ground-Truth Depth

DGX agent

arXiv:2605.19797v1 Announce Type: new Abstract: Monocular depth estimation has improved significantly in recent years, driven by increasingly powerful models and large-scale training data. Predicted d

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Detecting Fluent Optimization-Based Adversarial Prompts via Sequential Entropy Changes

DGX agent

arXiv:2605.19966v1 Announce Type: cross Abstract: Optimization-based adversarial suffixes can jailbreak aligned large language models (LLMs) while remaining fluent, weakening static and windowed perpl

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Did we ever learn what model won gold at the IMO from OpenAI? It was a year ago and it was called an unreleased internal general purpose mod…

DGX agent

Did we ever learn what model won gold at the IMO from OpenAI? It was a year ago and it was called an unreleased internal general purpose model back then. Has GPT-5.5 Pro Extended caught up with whatev

model-releasesethan-mollick--x
20 May 2026
Model Releases

Differential-Integral Neural Operator for Long-Term Turbulence Forecasting

DGX agent

arXiv:2509.21196v3 Announce Type: replace-cross Abstract: Accurately forecasting the long-term evolution of turbulence represents a grand challenge in scientific computing and is crucial for applicati

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Disentangling generalization and memorization in large language models using chess

DGX agent

arXiv:2601.16823v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) exhibit remarkable capabilities, yet it remains unclear to what extent these reflect sophisticated recall or genu

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization

DGX agent

arXiv:2605.18999v1 Announce Type: new Abstract: Muon and related normalized optimizers decouple the choice of update direction from the choice of step scale, but their practical performance remains se

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

Distilling Linearized Behavior for Effective Task Arithmetic

DGX agent

arXiv:2605.18993v1 Announce Type: cross Abstract: Task vector composition has emerged as a promising paradigm for editing pre-trained models, enabling model merging through addition and unlearning thr

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Distribution-Free Uncertainty Quantification for Continuous AI Agent Evaluation

DGX agent

arXiv:2605.19779v1 Announce Type: new Abstract: We adapt split conformal prediction and adaptive conformal inference (ACI) to continuous AI agent evaluation, providing distribution-free coverage guara

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Distributional Energy-Based Models for Uncertainty-Aware Structured LLM Reasoning

DGX agent

arXiv:2605.18871v1 Announce Type: cross Abstract: When Large Language Models produce structured outputs such as travel plans, code solutions, or multi-step proofs, individual reasoning steps may appea

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Distributionally Robust Control via Stein Variational Inference for Contact-Rich Manipulation

DGX agent

arXiv:2605.19029v1 Announce Type: new Abstract: Reliable robotic manipulation requires control policies that can accurately represent and adapt to uncertainty arising from contact-rich interactions. M

model-releasesarxiv-cs-ro
20 May 2026
Model Releases

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model

DGX agent

arXiv:2602.23622v2 Announce Type: replace-cross Abstract: Significant progress has been made in the field of Instruction-based Image Editing Models (IIEMs). However, while these models demonstrate pla

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

DGX agent

arXiv:2605.18915v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

DocQT: Improving Document Forgery Localization Robustness via Diverse JPEG Quantization Tables

DGX agent

arXiv:2605.19688v1 Announce Type: new Abstract: Document manipulation localization models achieve strong performance on public benchmarks yet fail to generalize to operational document workflows. We i

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Does Code Cleanliness Affect Coding Agents? A Controlled Minimal-Pair Study

DGX agent

arXiv:2605.20049v1 Announce Type: cross Abstract: As autonomous coding agents see rapid adoption, their evaluation has primarily focused on task completion rates holding the target codebase fixed. Thi

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Dual-Prompt CLIP with Hybrid Visual Encoders for Occluded Person Re-Identification

DGX agent

arXiv:2605.19527v1 Announce Type: new Abstract: Occluded person re-identification focuses on matching partially visible pedestrians across multiple camera views. However, occlusions disrupt body-regio

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

DualView: Adaptive Local-Global Fusion for Multi-Hop Document Reranking

DGX agent

arXiv:2605.18767v1 Announce Type: cross Abstract: Multi-hop question answering requires aggregating information from multiple documents, a critical capability for knowledge-intensive applications. A f

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Dynamic Model Merging Made Slim

DGX agent

arXiv:2605.18904v1 Announce Type: cross Abstract: Model merging enables the reuse of fine-tuned models without joint training or access to original data. Dynamic merging further improves flexibility b

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

DynaTrain: Fast Online Parallelism Switching for Elastic LLM Training

DGX agent

arXiv:2605.18815v1 Announce Type: new Abstract: Modern large language model (LLM) training is inherently dynamic: resource fluctuations, RLHF phase shifts, and cluster elasticity continually reshape t

model-releasesarxiv-cs-lg
20 May 2026
Model Releases

EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data

DGX agent

arXiv:2605.19130v1 Announce Type: cross Abstract: Children acquire language grounding with remarkable robustness from limited visuo-linguistic input in ways that surpass today's best large multimodal

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

DGX agent

arXiv:2605.19559v1 Announce Type: cross Abstract: The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the abil

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

EgoTraj: Real-World Egocentric Human Trajectory Dataset for Multimodal Prediction

DGX agent

arXiv:2605.19004v1 Announce Type: new Abstract: Accurately forecasting human trajectories from an egocentric perspective plays a central role in applications such as humanoid robotics, wearable sensin

model-releasesarxiv-cs-cv
20 May 2026
Model Releases

Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance

DGX agent

arXiv:2512.23461v2 Announce Type: replace-cross Abstract: Reward models (RMs) are essential in reinforcement learning from human feedback (RLHF) to align large language models (LLMs) with human values

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Embedding by Elicitation: Dynamic Representations for Bayesian Optimization of System Prompts

DGX agent

arXiv:2605.19093v1 Announce Type: new Abstract: System prompts are a central control mechanism in modern AI systems, shaping behavior across conversations, tasks, and user populations. Yet they are di

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Emergence of Frontier Superposition: Mobius attractor and Cascade Supervision

DGX agent

arXiv:2605.18820v1 Announce Type: cross Abstract: Superposition allows Transformers to reason in depth, carrying an entire reasoning frontier in parallel through a bounded-depth forward pass instead o

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

Enabling Real-Time Colonoscopic Polyp Segmentation on Commodity CPUs via Ultra-Lightweight Architecture

DGX agent

arXiv:2602.04381v2 Announce Type: replace-cross Abstract: Real-time polyp segmentation is essential for early colorectal cancer detection, yet clinical deployment remains blocked by GPU dependency. We

model-releasesarxiv-cs-ai
20 May 2026
Model Releases

EngiAI: A Multi-Agent Framework and Benchmark Suite for LLM-Driven Engineering Design

DGX agent

arXiv:2605.19743v1 Announce Type: new Abstract: Large Language Model (LLM) agents are increasingly applied to engineering design tasks, yet existing evaluation frameworks do not adequately address mul

model-releasesarxiv-cs-ai
20 May 2026
← Previous
1…284285286287288…472
Next →