AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

88,419Total entries
1Added by human
88,418Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,638 results
26 May 2026

Geometric Evolution Maps: Extracting Stable Concept Probes from Transformer Residual Streams

ResearchDGX agent

arXiv:2605.25848v1 Announce Type: cross Abstract: Concept probes extracted from transformer residual streams are only as reliable as the layer from which they are extracted. The common practice of pro

GeoSVG-RL: Geometry-Aware Reinforcement Learning for Layout-Constrained Text-to-SVG Diagram Generation

SafetyDGX agent

arXiv:2605.25447v1 Announce Type: new Abstract: Generating structured, editable diagrams remains a significant challenge for contemporary large language models, despite their proficiency in general-pu

GroupTravelBench: Benchmarking LLM Agents on Multi-Person Travel Planning

Model ReleasesDGX agent

arXiv:2605.25200v1 Announce Type: new Abstract: Travel planning is a realistic task for evaluating the planning and tool-use abilities of LLM agents. However, existing benchmarks typically assume only

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Guided Flow Matching for Forward and Inverse PDE Problems with Sparse Observations: Algorithm and Theory

Model ReleasesDGX agent

arXiv:2605.25509v1 Announce Type: cross Abstract: Reconstructing PDE solutions from sparse observations is a core challenge in scientific computing. We present FM4PDE, a flow-matching generative frame

Hera: Learning Long-Horizon Coordination for Device-Cloud Collaborative LLM Agents

Local AiDGX agent

arXiv:2605.24598v1 Announce Type: new Abstract: Large language model (LLM) agents excel at solving complex long-horizon tasks through autonomous interaction with environments. However, their real-worl

HiGraph: A Large-Scale Hierarchical Graph Dataset for Malware Analysis

Model ReleasesDGX agent

arXiv:2509.02113v2 Announce Type: replace-cross Abstract: The advancement of graph-based malware analysis is critically limited by the absence of large-scale datasets that capture the inherent hierarc

How Many Tools Should an LLM Agent See? A Chance-Corrected Answer

Model ReleasesDGX agent

arXiv:2605.24660v1 Announce Type: cross Abstract: Before an LLM agent can use a tool, a retrieval system must decide which candidate tools to show to the agent. How long should that shortlist be? Show

I found this Wired article on AI fact-checking frustrating. It could have been about why we continue to need human fact checkers (talk to pe…

Model ReleasesDGX agent

I found this Wired article on AI fact-checking frustrating. It could have been about why we continue to need human fact checkers (talk to people, use judgement, resolve conflict). Instead it is full o

IterInject: Indirect Prompt Injection Against LLM Agents via Feedback-Guided Iterative Optimization

Model ReleasesDGX agent

arXiv:2605.24659v1 Announce Type: new Abstract: LLM-based agents are increasingly deployed for complex tasks requiring planning, tool use, and interaction with external services. Their reliance on unt

Keep the Proof State Live: Snapshotting for Efficient Tactic Search in Lean 4

Model ReleasesDGX agent

arXiv:2605.25556v1 Announce Type: cross Abstract: Automated theorem proving systems built on Lean 4 increasingly rely on parallel tactic search over partially specified proofs, such as those generated

Kolmogorov-Arnold Fourier Networks

Model ReleasesDGX agent

arXiv:2502.06018v3 Announce Type: replace-cross Abstract: Although Kolmogorov-Arnold-based interpretable networks (KANs) possess strong theoretical expressiveness, they suffer from severe parameter ex

LLMs on prem. but the prem is your wardrobe. http://shop.cohere.com The design team cooked on this one.

Model ReleasesDGX agent

Cohere has announced a creative marketing initiative featuring LLM-themed merchandise available through their shop, with the tagline playing on the phrase 'on premises' to humorously suggest their lan

Lngram: N-gram Conditional Memory in Latent Space

Local AiDGX agent

arXiv:2605.24869v1 Announce Type: new Abstract: Sequence modeling requires both compositional reasoning and local static knowledge retrieval, yet standard Transformers handle both through dense comput

LWM-CDE: A Representation Space for Wireless Data Reasoning and Transferability

ApplicationsDGX agent

arXiv:2605.24077v1 Announce Type: cross Abstract: Machine learning deployments in real-world wireless communication tasks face significant generalization challenges due to location and environment-spe

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework

Model ReleasesDGX agent

arXiv:2605.24661v1 Announce Type: new Abstract: LLMs have achieved remarkable success in complex reasoning tasks, yet current evaluation approaches predominantly rely on final-answer correctness, offe

Memory-Induced Tool-Drift in LLM Agents

Model ReleasesDGX agent

arXiv:2605.24941v1 Announce Type: cross Abstract: Modern LLM agents combine long-term memory for personalization with tool-calling interfaces for taking actions in the world -- a combination underpinn

MIND: Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control

Model ReleasesDGX agent

arXiv:2605.26006v1 Announce Type: cross Abstract: Enabling physics-based humanoids to execute diverse behaviors from high-level textual commands remains a significant challenge. Existing methods typic

MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding

Model ReleasesDGX agent

arXiv:2605.24523v1 Announce Type: cross Abstract: Visual decoding from brain signals is a key challenge at the intersection of computer vision and neuroscience, requiring methods that bridge neural re

Multitask learning with semiempirical orbital charges enables sample-efficient MLIPs

TutorialsDGX agent

arXiv:2605.24073v1 Announce Type: cross Abstract: Machine learning interatomic potentials (MLIPs) require generating computationally expensive, large-scale training datasets to accurately simulate mat

Muon in Associative Memory Learning: Training Dynamics and Scaling Laws

Model ReleasesDGX agent

arXiv:2602.05725v2 Announce Type: replace Abstract: Muon updates matrix parameters via the matrix sign of the gradient and has shown strong empirical gains, yet its dynamics and scaling behavior remai

Nine reasons why I warned OpenAI might fail and turn out to be the WeWork of AI, from two years ago.

Model ReleasesDGX agent

Nine reasons why I warned OpenAI might fail and turn out to be the WeWork of AI, from two years ago. 9 reasons that OpenAI could someday be seen as the WeWork of AI: 👉 Lots of competitors are catching

Noise-Robust Financial Numerical Entity Attribute Tagging

Model ReleasesDGX agent

arXiv:2605.24910v1 Announce Type: new Abstract: Financial Numerical Entity (FNE) understanding aims to recover the meaning of numerical mentions in financial reports. Existing studies primarily focus

NormimesDirection: Restoring the Missing Query Norm in Vision Linear Attention

Model ReleasesDGX agent

arXiv:2506.21137v3 Announce Type: replace Abstract: Linear attention mitigates the quadratic complexity of softmax attention but suffers from a critical loss of expressiveness. We identify two primary

On the Epistemic Uncertainty of Overparametrized Neural Networks

Model ReleasesDGX agent

arXiv:2605.25234v1 Announce Type: cross Abstract: Epistemic uncertainty is often viewed as a reducible uncertainty that vanishes with increasing data. This perspective implicitly assumes parameter ide

Parallel Differentiable Reachability for Learning and Planning with Certified Neural Dynamics and Controllers

HardwareDGX agent

arXiv:2605.25346v1 Announce Type: cross Abstract: Neural network (NN) dynamics models and control policies achieve strong performance in robotics, but providing sound guarantees under uncertainty rema

Partner-Aware Hierarchical Skill Discovery for Robust Human-AI Collaboration

Model ReleasesDGX agent

arXiv:2605.24352v1 Announce Type: new Abstract: Multi-agent collaboration, especially in human-AI teaming, requires agents that can adapt to novel partners with diverse and dynamic behaviors. Conventi

{Phi}-Noise: Training-Free Temporal Video Conditioning via Phase-Based Noise Manipulation

ResearchDGX agent

arXiv:2605.24509v1 Announce Type: cross Abstract: Latent video diffusion models generate videos by progressively transforming Gaussian noise into realistic samples conditioned on text or visual inputs

Poisoning the Watchtower: Prompt Injection Attacks Against LLM-Augmented Security Operations Through Adversarial Log Content

ResearchDGX agent

arXiv:2605.24421v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as analyst assistants in security operations centers (SOCs), where they ingest log and alert data t

Position: AI for Science Should Treat Measurement-to-Dataset Pipelines as Inference Components

Model ReleasesDGX agent

arXiv:2605.24558v1 Announce Type: new Abstract: AI for Science (AI4Science) workflows often treat the released dataset as a fixed interface to the underlying system. However, in domains relying on ind

PowLU: An Activation Function for Stable Pre-Training of LLMs

ResearchDGX agent

arXiv:2605.25704v1 Announce Type: new Abstract: In contemporary large language models (LLMs), the swish-gated linear unit (SwiGLU) activation function is widely adopted to regulate the information flo

Prefix Teach, Suffix Fade: Local Teachability Collapse in Strong-to-Weak On-Policy Distillation

Local AiDGX agent

arXiv:2605.13643v2 Announce Type: replace Abstract: On-policy distillation (OPD) trains a student model on its own rollouts using dense feedback from a stronger teacher. Prior literature suggests that

Querying structural and functional niches on spatial transcriptomics data

Model ReleasesDGX agent

arXiv:2410.10652v4 Announce Type: replace-cross Abstract: Cells in multicellular organisms coordinate to form structural and functional niches. With spatial transcriptomics (ST) enabling gene expressi

Reasoning as an Attack Surface: Adaptive Evolutionary CoT Jailbreaks for LLMs

SafetyDGX agent

arXiv:2605.24497v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) have demonstrated remarkable capabilities in reasoning and generation tasks and are increasingly deployed in real-world ap

RED: Adaptive Real-Time DAG Scheduling for Robotic Inference under Environmental Dynamics

Model ReleasesDGX agent

arXiv:2605.24044v1 Announce Type: new Abstract: Robots deployed in dynamic environments must contend with environment-driven changes that reshape computation at runtime: new tasks may appear, preceden

Reinforcement Learning from Denoising Feedback

SafetyDGX agent

arXiv:2605.25638v1 Announce Type: new Abstract: Policy loss estimation remains a fundamental and long-standing challenge in reinforcement learning (RL) for diffusion language models (dLLMs). We introd

Remote sensing data imputation using deep learning for multispectral imagery

ResearchDGX agent

arXiv:2605.24003v1 Announce Type: cross Abstract: Remote sensing techniques have been increasingly utilised in aquatic applications in recent years. A common challenge in using optical satellite data

Rewarding Structural Conformance of Reasoning using Process Mining

SafetyDGX agent

arXiv:2510.25065v3 Announce Type: replace Abstract: Recent advances in sparse reward policy gradient methods have enabled effective reinforcement learning (RL)-based language model post-training. Howe

Safety Generalization Under Distribution Shift in Safe Reinforcement Learning: A Diabetes Testbed

Model ReleasesDGX agent

arXiv:2601.21094v2 Announce Type: replace-cross Abstract: Safe Reinforcement Learning (RL) algorithms are typically evaluated under fixed training conditions. We investigate whether training-time safe

SAM: State-Adaptive Memory for Long-Horizon Reasoning Agent

AgentsDGX agent

arXiv:2605.24468v1 Announce Type: new Abstract: Long-horizon agentic reasoning requires large language models to act over long interaction histories containing thoughts, tool calls, observations, and

SentGraph: Hierarchical Sentence Graph for Multi-hop Retrieval-Augmented Question Answering

ResearchDGX agent

arXiv:2601.03014v3 Announce Type: replace-cross Abstract: Traditional Retrieval-Augmented Generation (RAG) effectively supports single-hop question answering with large language models but faces signi

'Si'multaneous 'S'patial-'T'emporal Message Passing for Dynamic Graph Representation Learning

Model ReleasesDGX agent

arXiv:2605.25548v1 Announce Type: cross Abstract: Dynamic graph neural networks (DGNNs) that operate on snapshot sequences typically fall into one of two categories. Temporal-first approaches build pe

Smart Timing for Mining: A Deep Learning Framework for Bitcoin Hardware ROI Prediction

Model ReleasesDGX agent

arXiv:2512.05402v2 Announce Type: replace-cross Abstract: Bitcoin mining hardware acquisition requires strategic timing due to volatile markets, rapid technological obsolescence, and protocol-driven r

SPARK: Search Personalization via Agent-Driven Retrieval and Knowledge-sharing

AgentsDGX agent

arXiv:2512.24008v3 Announce Type: replace Abstract: Personalized search demands the ability to model users' evolving, multi-dimensional information needs; a challenge for systems constrained by static

Spiking the training data to correct for test set contamination

ResearchDGX agent

arXiv:2605.24818v1 Announce Type: cross Abstract: The literature on test set contamination largely focuses on detection, but the correction of contaminated test scores is underexplored. Our core propo

StakeBench: Evaluating Language Understanding Grounded in Market Commitment

SafetyDGX agent

arXiv:2605.26074v1 Announce Type: cross Abstract: Existing financial NLP benchmarks often rely on labels supplied by outside observers, measuring how language is perceived rather than what speakers ha

Temporal Concept Drift in Legal Judgment Prediction: Neural Baselines Across Three Epochs of Ukrainian Court Decisions

ApplicationsDGX agent

arXiv:2605.24452v1 Announce Type: cross Abstract: Legal NLP benchmarks evaluate models on randomly split data, implicitly assuming that legal language is stationary. We test this assumption by fine-tu

The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks

SafetyDGX agent

arXiv:2602.16340v3 Announce Type: replace Abstract: We study the implicit bias of momentum-based optimizers on smooth homogeneous models. We show that extit{momentum steepest descent} algorithms like

this logic from 2024 held up pretty well considering all that’s changed.

Model ReleasesDGX agent

this logic from 2024 held up pretty well considering all that’s changed. 9 reasons that OpenAI could someday be seen as the WeWork of AI: 👉 Lots of competitors are catching up. 👉 OpenAI has been force

TinyFormer: Preserving Tiny Objects in YOLO-DETRHybridReal-time Detectors

ResearchDGX agent

arXiv:2605.25046v1 Announce Type: cross Abstract: YOLO-series and DETR-based detectors struggle with tiny-object detection. YOLO-style models benefit from efficient dense prediction, but their large-s

TSFLora: Token-Compressed Split Fine-Tuning for Wireless Edge Networks

ResearchDGX agent

arXiv:2605.23988v1 Announce Type: cross Abstract: Adapting large AI models (LAMs) to personalized edge data is challenging because wireless devices have limited memory, computation, and uplink capacit

TTPrint: Evidence-Grounded TTP Extraction via Diverge-then-Converge Verification

Model ReleasesDGX agent

arXiv:2605.25836v1 Announce Type: cross Abstract: Extracting MITRE ATT&CK techniques from cyber threat intelligence (CTI) reports is an open-set, multi-label problem requiring both high recall (not mi

Turn-Based Structural Triggers: Prompt-Free Backdoors in Multi-Turn LLMs

ResearchDGX agent

arXiv:2601.14340v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are widely integrated into interactive systems such as dialogue agents and task-oriented assistants. This growing

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification

Model ReleasesDGX agent

arXiv:2605.25474v1 Announce Type: new Abstract: TypedCSIP is a typed counterfactual pretraining method for the conflict-classification task of the LCR-CN benchmark (Zhao et al., 2026): given a (superi

v0.30.0-rc26: Merge remote-tracking branch 'upstream/main' into llama-runner-phase-0

Model ReleasesDGX agent

This release candidate merges updates from the upstream main branch into the llama-runner-phase-0 branch, likely incorporating recent improvements and bug fixes into the development version. Version 0

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

Model ReleasesDGX agent

arXiv:2605.25988v1 Announce Type: new Abstract: Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. extbf{We find that the check

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

Model ReleasesDGX agent

arXiv:2605.23932v1 Announce Type: new Abstract: Despite strong medical benchmark accuracy, LLMs can exhibit severe multi-turn sycophancy in clinical dialogue, abandoning initial correct diagnosis unde

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

Model ReleasesDGX agent

arXiv:2605.24902v1 Announce Type: cross Abstract: Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical do

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards

SafetyDGX agent

arXiv:2605.25864v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable advancements in reasoning capabilities empowered by Reinforcement Learning with Verifiable Rewar

WLNO: Wavelet-Laplace Neural Operator for Solving Partial Differential Equations

Model ReleasesDGX agent

arXiv:2605.24658v1 Announce Type: new Abstract: This work introduces the Wavelet-Laplace Neural Operator (WLNO), a novel neural operator that fuses Haar wavelet multi-scale spatial decomposition with

WorldGUI: An Interactive Benchmark for Desktop GUI Automation from Any Starting Point

Model ReleasesDGX agent

arXiv:2502.08047v5 Announce Type: replace Abstract: Recent progress in GUI agents has substantially improved visual grounding, yet robust planning remains challenging, particularly when the environmen

← Previous
1…549550551552553…1061
Next →