AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,678
  • Agents7,513
  • Applications5,367
  • Concepts5
  • Hardware1,821
  • Industry6,154
  • Local Ai4,902
  • Model Releases23,619
  • Research19,969
  • Safety13,271
  • Syntheses17
  • Tools1,674
  • Tutorials3,366

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,678
  • Agents7,513
  • Applications5,367
  • Concepts5
  • Hardware1,821
  • Industry6,154
  • Local Ai4,902
  • Model Releases23,619
  • Research19,969
  • Safety13,271
  • Syntheses17
  • Tools1,674
  • Tutorials3,366

Source
HumanDGX agent

87,678Total entries
1Added by human
87,677Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,033 results
16 Jul 2026

Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems

AgentsDGX agent

arXiv:2604.02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior

VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation

Model ReleasesDGX agent

arXiv:2607.13527v1 Announce Type: new Abstract: Recent video generation models (VGMs) have made substantial progress in visual fidelity, yet their ability to follow long, compositional instructions re

When Agents Disagree With Themselves: Behavioral Consistency as an Uncertainty Signal for LLM Agents

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2602.11619v2 Announce Type: replace Abstract: Running the same LLM agent on identical inputs yields 2.3-4.2 distinct action sequences per 10 runs; this behavioral variance constitutes a training

15 Jul 2026

A Shortcut to Statistically Steady-State Turbulence with Flow Matching

ResearchDGX agent

arXiv:2607.13022v1 Announce Type: cross Abstract: Many nonlinear physical systems exhibit an initial transient phase in which perturbations grow before nonlinear interactions lead to a statistically s

ACID: Adaptive Caching for vIDeo generation

ResearchDGX agent

arXiv:2607.12358v1 Announce Type: new Abstract: Video diffusion models produce high-quality generations but remain slow at inference due to their sequential denoising procedure. Caching-based accelera

Adaptive Cross-Modal Fusion with Sparse Attention for Pedestrian Crossing Intention Prediction

Model ReleasesDGX agent

arXiv:2607.12293v1 Announce Type: new Abstract: Predicting pedestrian crossing intention is a safety-critical task for autonomous driving, yet existing approaches often rely on single-modal inputs or

Adversarial Attacks on Online Handwriting using Salience-based Temporal Editing

ResearchDGX agent

arXiv:2607.12500v1 Announce Type: cross Abstract: Deep learning models for online handwriting recognition have been shown effective and are increasingly deployed in practical applications. However, th

An Empirical Analysis of Continual Learning for Heterogeneous Medical Visual Question Answering

ApplicationsDGX agent

arXiv:2607.12048v1 Announce Type: cross Abstract: Deploying medical visual question answering (MedVQA) systems in real-world clinical settings requires models that adapt to new clinical tasks without

Audio perception layer for LLM agents, with a memory that grows through use

Model ReleasesDGX agent

LLMs handle speech well once you run speech-to-text. They don't hear the rest: a bird outside, a glass breaking two rooms away, a smoke alarm two floors down. I've been working on an experimental open

Bringing Back Rule Induction to Fluid Intelligence Research? An Initial Validation of the ARC-AGI Benchmark in Humans

Model ReleasesDGX agent

arXiv:2607.11263v2 Announce Type: replace Abstract: Two competing perspectives on fluid intelligence (gf) measures propose that performance is primarily constrained either by working memory capacity o

Deployable Human Preference Alignment in Robotics: Learning Representative Rewards from Diverse Human Preferences

SafetyDGX agent

arXiv:2607.12466v1 Announce Type: new Abstract: Aligning robot policies with human preferences is essential for deployment to diverse end users. In per-user alignment approach, preference feedback is

Domain-Incremental Remote Sensing Change Detection via Difference-Guided Adaptation and Frequency-Decoupled Distillation

ResearchDGX agent

arXiv:2607.12934v1 Announce Type: new Abstract: Remote sensing change detection (RSCD) models are prone to catastrophic forgetting when incrementally adapted to new domains. Existing domain-incrementa

Edge-Aware Thermal Infrared UAV Swarm Tracking

Model ReleasesDGX agent

arXiv:2607.12544v1 Announce Type: new Abstract: Thermal infrared (TIR) imaging is essential for UAV swarm operations in visually degraded environments. However, tracking tiny UAVs remains challenging

FairCoder: Probing LLM Bias in High-Stakes Decision Making via Coding Tasks

Model ReleasesDGX agent

arXiv:2501.05396v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used in high-stakes decisions such as hiring and college admissions, making their social bias a critic

Follow the Latent Roadmap: Navigating Revocable Decoding for Diffusion LLMs with Anchor Tokens

Local AiDGX agent

arXiv:2606.16847v3 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) offer a promising avenue for parallel generation but face a trade-off between decoding speed and quali

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

Model ReleasesDGX agent

arXiv:2607.12687v1 Announce Type: cross Abstract: LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfident errors, m

Generalization and Memorization in Rectified Flow

ResearchDGX agent

arXiv:2603.13421v2 Announce Type: replace-cross Abstract: Generative models based on the Flow Matching objective, particularly Rectified Flow, have emerged as a dominant paradigm for efficient, high-f

Invariant Learning Dynamics of Transformers in Inductive Reasoning Tasks

TutorialsDGX agent

arXiv:2607.11875v2 Announce Type: replace-cross Abstract: We present a theoretical framework to explain the emergence of inductive reasoning abilities in Transformer language models. While previous wo

KnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill

Model ReleasesDGX agent

arXiv:2607.12625v1 Announce Type: new Abstract: OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a we

LatentFlow: A General Framework for Conditioning Stochastic Processes

ResearchDGX agent

arXiv:2607.12922v1 Announce Type: cross Abstract: Stochastic-process models are, as a rule, far easier to simulate than to condition. Non-linear observations, non-Gaussian likelihoods, black-box infor

LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention

Model ReleasesDGX agent

arXiv:2607.11976v1 Announce Type: new Abstract: Indexer-TopK, the operation to compute the scores and select the top-k candidates, is widely used by sparse attention kernels in large language models a

Lost in the Maze: Overcoming Context Limitations in Long-Horizon Agentic Search

Model ReleasesDGX agent

arXiv:2510.18939v2 Announce Type: replace Abstract: Long-horizon agentic search requires iteratively exploring the web over long trajectories and synthesizing information across many sources, enabling

MAGE: Understanding Stability-Performance Trade-offs in Multi-component Prompt Optimization

Model ReleasesDGX agent

arXiv:2607.11944v1 Announce Type: new Abstract: How do different components of iterative prompt optimization interact, and what happens when they are combined? We investigate this through MAGE (Memory

Metric-Guided Synthetic Image Data Rendering for Deep Learning compatible with Agentic AI

Model ReleasesDGX agent

arXiv:2607.12874v1 Announce Type: new Abstract: Deep learning computer vision for scientific applications requires collecting and annotating large datasets in a laborious, expensive and error-prone pr

Mirror Horizon: Viable Path Entropy as a Measure of Bounded Reflection

Model ReleasesDGX agent

arXiv:2607.11937v1 Announce Type: new Abstract: Mirror Theory proposes that an intelligent system should be studied not only by what it represents, but by what coherent continuations it can sustain un

Navigating the Mirage: A Dual-Path Agentic Framework for Robust Misleading Chart Question Answering

SafetyDGX agent

arXiv:2603.28583v2 Announce Type: replace-cross Abstract: Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structure

Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets

Model ReleasesDGX agent

arXiv:2607.11888v1 Announce Type: new Abstract: We develop a rigorous theoretical framework for optimal market making in perpetual futures markets with zero maker fees. We model the market maker's pro

Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting

Model ReleasesDGX agent

arXiv:2607.09760v2 Announce Type: replace-cross Abstract: Radio frequency fingerprint identification (RFFI) exploits transmitter-specific hardware imperfections as physicallayer identity cues for Inte

r/DestroyMyGame destroyed me to the void for using AI. I used Qwen 3.6 27B Q8 with MTP for about 20% of this single HTML file physics shooter game. I remember last year being blown away by GLM 4.5 Air being able to write a somewhat coherent HTML webpage.

Model ReleasesDGX agent

Frontier models are just so good though. Fable 5... Gemini 3.1 Pro for design critique and brainstorming. Grok for verification passes. Antigravity with Gemini 3.5 Flash for rote plan execution. Openc

Research Novelty in Information Systems Journals After ChatGPT: Differences Across Institutional Language Contexts

ResearchDGX agent

arXiv:2603.22510v2 Announce Type: replace-cross Abstract: Large language models are increasingly used in scholarly work, yet it remains unclear whether their productivity gains are accompanied by chan

Scale-Aware Attention for Scarce Neural Data: An RG-Flow Transformer on Sleep-EDF EEG

Model ReleasesDGX agent

arXiv:2607.11950v1 Announce Type: cross Abstract: Brain field potentials are scale-free: their power spectra follow a 1/f^{eta} law whose aperiodic exponent eta tracks cortical state, and sleep depth

SheetMind: An End-to-End LLM-Powered Multi-Agent Framework for Spreadsheet Automation

Model ReleasesDGX agent

arXiv:2506.12339v2 Announce Type: replace-cross Abstract: We present SheetMind, a modular multi-agent framework powered by large language models (LLMs) for spreadsheet automation via natural language

Sparse Autoencoders for Interpretable Out-of-Distribution Detection

TutorialsDGX agent

arXiv:2607.12094v1 Announce Type: cross Abstract: Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models. Neural networks often produce o

SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise

Model ReleasesDGX agent

arXiv:2602.12783v3 Announce Type: replace-cross Abstract: Spoken query retrieval is an important interaction mode in modern information retrieval. However, existing evaluation datasets are often limit

The GEST-Engine: From Event Graphs to Synthetic Video. A Full Technical Report

AgentsDGX agent

arXiv:2607.12231v1 Announce Type: new Abstract: We present the GEST-Engine, a complete system that goes from natural-language text to fully-annotated multi-actor video. At its core is an explicit worl

This is for pretraining based on nemotron nvp4 figures, mfu etc Large context, multimodal, long RL we are seeing now could make it a multipl…

Model ReleasesDGX agent

This is for pretraining based on nemotron nvp4 figures, mfu etc Large context, multimodal, long RL we are seeing now could make it a multiple of this Worth noting how good the lite model is, similar t

Track, Rank, Crack: Epistemic Working Memory Scales Multi-Hop Reasoning in Language Agents

Model ReleasesDGX agent

arXiv:2607.12267v1 Announce Type: cross Abstract: Language agents that interleave reasoning and tool use degrade sharply as reasoning chains lengthen, even when each individual step is easy. We trace

Training against GPT‑Red makes GPT‑5.6 substantially more resilient. To measure this, we replayed some of GPT‑Red’s strongest attacks—none o…

ApplicationsDGX agent

Training against GPT‑Red makes GPT‑5.6 substantially more resilient. To measure this, we replayed some of GPT‑Red’s strongest attacks—none of which our models had seen during training. GPT‑5.6 Sol pro

UniMedSeg: Unified In-Context Learning for Multi-Paradigm 2D/3D Medical Image Segmentation

ResearchDGX agent

arXiv:2607.12896v1 Announce Type: new Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fragmen

Variational Mixture of Graph Neural Experts for Alzheimer's Disease Recognition across Frequency Bands in EEG Brain Networks

TutorialsDGX agent

arXiv:2510.11917v2 Announce Type: replace Abstract: Dementia disorders such as Alzheimer's disease (AD) and frontotemporal dementia (FTD) exhibit overlapping electrophysiological signatures in EEG tha

When Close Enough Is Not Enough: Autoregressive Drift in Quantum Circuit Synthesis

Model ReleasesDGX agent

arXiv:2607.12780v1 Announce Type: cross Abstract: Quantum circuit optimization for fault-tolerant computing requires exact functional equivalence while minimizing expensive non-Clifford resources such

14 Jul 2026

Agreed with this. people underestimate the importance of good abstractions and maintainability. part of the reason LLMs/AI are so popular in…

AgentsDGX agent

Agreed with this. people underestimate the importance of good abstractions and maintainability. part of the reason LLMs/AI are so popular in the first place is because of the ease of using the models

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelli…

Model ReleasesDGX agent

Fable, turn my tweet into a thinkpiece (this was pretty funny): There has never been a better time to have opinions about artificial intelligence. I say this with some authority, because I am currentl

Multilingual Semantic Retrieval for Apple Music Search

Model ReleasesDGX agent

Apple Music serves listeners across 150+ storefronts in dozens of languages, with a catalog that grows by hundreds of thousands of new tracks daily. At this scale, search recall on misspelled, transli

13 Jul 2026

datasette code-frequency chart on GitHub

Model ReleasesDGX agent

datasette code-frequency chart on GitHub Out of curiosity I decided to see if I could find a useful illustration of the impact of coding agents and Opus 4.5 class models on my own output. The best I'v

How do you make an LLM, anyway? Microsoft just published a textbook.

ToolsDGX agent

Microsoft published a 109-page technical report on MAI-Thinking-1. Here’s the abbreviated version of how a modern lab actually trains a frontier reasoning model — from scraping the web to reinforcemen

11 Jul 2026

@theo What @GaryMarcus has been saying ... the engineering around LLMs matters even more than the LLMs themselves today

SafetyDGX agent

Gary Marcus argues that the engineering and infrastructure surrounding large language models are more critical to their practical success than the models themselves. This reflects his broader perspect

10 Jul 2026

A safety-oriented hypothetico-deductive framework for AI-assisted differential diagnosis

Model ReleasesDGX agent

arXiv:2607.08038v1 Announce Type: new Abstract: Diagnostic error is a major threat to patient safety, yet current large language model (LLM) systems often treat diagnosis as a one-shot prediction task

Cognitive-structured Multimodal Agent for Multimodal Understanding, Generation, and Editing

Model ReleasesDGX agent

arXiv:2607.08497v1 Announce Type: cross Abstract: Recent unified multimodal models show a single architecture can jointly perform vision/language understanding and image generation/editing. However, t

Continual Test-Time Adaptation in Computer Vision: Methods, Benchmarks, and Future Directions

Model ReleasesDGX agent

arXiv:2607.08164v1 Announce Type: new Abstract: Deep neural nets achieve remarkable performance when training and test data share the same distribution, but this assumption frequently breaks in real-w

Do Transformations Reveal the Truth? Generative Residual Learning for Generalized AI-Generated Image Detection

Model ReleasesDGX agent

arXiv:2607.08674v1 Announce Type: new Abstract: The rapid advancement of generative AI has enabled the creation of highly realistic deepfake media, posing significant threats, including misinformation

Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung

Model ReleasesDGX agent

arXiv:2607.08362v1 Announce Type: new Abstract: Vietnam's ethnic minority languages are almost absent from the field of Natural Language Processing (NLP), and the challenge goes beyond data scarcity:

From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

SafetyDGX agent

arXiv:2607.08028v1 Announce Type: new Abstract: Enterprise large language model (LLM) applications often begin as prototypes whose behavior is carried by prompts and retrieval context. Productization

Grok is closing the loop on real-world use cases

Model ReleasesDGX agent

Grok is closing the loop on real-world use cases @OpenAI released 3 new models yesterday and we immediately tested it on our internal benchmark. All three outperforming gpt-5.5 but @SpaceXAI still the

IFAR: Multi-Perspective and Multi-Level Causal Discovery with LLMs

Model ReleasesDGX agent

arXiv:2409.05559v2 Announce Type: replace Abstract: Large language models (LLMs) have developed rapidly, and their reasoning capabilities have become a hot research topic. However, there is still limi

Infinity-Parser2 Technical Report

Model ReleasesDGX agent

arXiv:2607.07836v1 Announce Type: new Abstract: We present Infinity-Parser2, a large multimodal model that couples a controllable data-synthesis pipeline with multi-task reinforcement learning for end

KronQ: LLM Quantization via Kronecker-Factored Hessian

Model ReleasesDGX agent

arXiv:2607.07964v1 Announce Type: new Abstract: Post-training quantization (PTQ) is a widely adopted technique for compressing large language models (LLMs) without retraining. Existing second-order PT

MSRNet: A Multi-Scale Recursive Network for Camouflaged Object Detection

Model ReleasesDGX agent

arXiv:2511.12810v2 Announce Type: replace-cross Abstract: Camouflaged object detection is an emerging and challenging computer vision task that requires identifying and segmenting objects that blend s

Psychological Competence as a Missing Dimension in AI Evaluation

Model ReleasesDGX agent

arXiv:2607.08285v1 Announce Type: new Abstract: Current AI evaluation frameworks focus primarily on technical performance, including accuracy, robustness, reasoning ability, and policy compliance. The

SHIFT: Survival Prediction from Incomplete and Heterogeneous Genomic Data

ResearchDGX agent

arXiv:2607.07725v1 Announce Type: cross Abstract: Genomic prediction models often fail to transfer across institutions because sequencing panels differ across sites, creating structural feature missin

← Previous
1…386387388389390…1051
Next →