AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,082 results
12 May 2026

Not All Proofs Are Equal: Evaluating LLM Proof Quality Beyond Correctness

Model ReleasesDGX agent

arXiv:2605.10379v1 Announce Type: new Abstract: Large language models (LLMs) have become capable mathematical problem-solvers, often producing correct proofs for challenging problems. However, correct

Omni-DeepSearch: A Benchmark for Audio-Driven Omni-Modal Deep Search

Model ReleasesDGX agent

arXiv:2605.08762v1 Announce Type: cross Abstract: Current omni-modal benchmarks mainly evaluate models under settings where multiple modalities are provided simultaneously, while the ability to start

Overconfident and Blind to Details: Fixing Prompt Insensitivity with Abductive Preference Learning

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.09887v2 Announce Type: replace Abstract: Vision and language models frequently ignore semantically critical input edits, defaulting to pretraining priors. For example, models will confident

Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes

Model ReleasesDGX agent

arXiv:2602.00593v2 Announce Type: replace Abstract: Despite progress on general tasks, vision-language models (VLMs) still struggle with challenges that demand both fine-grained visual grounding and e

Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark

Model ReleasesDGX agent

arXiv:2410.14702v2 Announce Type: replace Abstract: Multi-modal Large Language Models (MLLMs) exhibit impressive problem-solving abilities in various domains, but their visual comprehension and abstra

ReorgGS: Equivalent Distribution Reorganization for 3D Gaussian Splatting

Model ReleasesDGX agent

arXiv:2605.08739v1 Announce Type: new Abstract: A converged 3D Gaussian Splatting (3DGS) model may approximate the target scene while remaining poorly parameterized for further optimization. We identi

Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs

Model ReleasesDGX agent

arXiv:2605.09239v1 Announce Type: new Abstract: Large language models fail at counting repeated tokens despite strong performance on broader reasoning benchmarks. These failures are commonly attribute

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

SafetyDGX agent

arXiv:2605.08257v1 Announce Type: cross Abstract: Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develo

Seed Hijacking of LLM Sampling and Quantum Random Number Defense

Model ReleasesDGX agent

arXiv:2605.08313v1 Announce Type: cross Abstract: Large language models (LLMs) rely on deterministic pseudorandom number generators (PRNGs) for autoregressive sampling, creating a critical supply-chai

SEMASIA: A Large-Scale Dataset of Semantically Structured Latent Representations

Model ReleasesDGX agent

arXiv:2605.09485v1 Announce Type: new Abstract: Latent representations learned by neural networks often exhibit semantic structure, where concept similarity is reflected by geometric proximity in embe

Sequential Membership Inference Attacks

ResearchDGX agent

arXiv:2602.16596v2 Announce Type: replace Abstract: Modern AI models are not static. They go through multiple updates in their lifecycles. We propose to design Sequential Membership Inference (SeMI) a

SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy

Model ReleasesDGX agent

arXiv:2605.09598v1 Announce Type: new Abstract: Vision-language models (VLMs) have recently shown strong potential in soccer video understanding. However, given the high complexity of soccer videos du

Soohak: A Mathematician-Curated Benchmark for Evaluating Research-level Math Capabilities of LLMs

Model ReleasesDGX agent

arXiv:2605.09063v1 Announce Type: new Abstract: Following the recent achievement of gold-medal performance on the IMO by frontier LLMs, the community is searching for the next meaningful and challengi

SPDEBench: An Extensive Benchmark for Learning Stochastic PDEs

Model ReleasesDGX agent

arXiv:2505.18511v2 Announce Type: replace Abstract: Stochastic Partial Differential Equations (SPDEs) driven by random noise play a central role in modeling physical processes with rough spatio-tempor

Statistical Scouting Finds Debate-Safe but Not Debate-Useful Cases: A Matched-Ceiling Study of Open-Weight LLM Reasoning Protocols

Model ReleasesDGX agent

arXiv:2605.09618v1 Announce Type: new Abstract: When should a language model answer directly, sample and vote, or engage in multi-agent debate? Recent work shows voting often explains much of the gain

Test-Time Training for Visual Foresight Vision-Language-Action Models

ResearchDGX agent

arXiv:2605.08215v1 Announce Type: new Abstract: Visual Foresight VLA (VF-VLA) has become a prominent architectural choice in the recent VLA due to its impressive performance. Nevertheless, the inheren

The Gordian Knot for VLMs: Diagrammatic Knot Reasoning as a Hard Benchmark

Model ReleasesDGX agent

arXiv:2605.09900v1 Announce Type: new Abstract: A vision-language model can look at a knot diagram and report what it sees, yet fail to act on that structure. KnotBench pairs an 858,318-image corpus f

Unmasking On-Policy Distillation: Where It Helps, Where It Hurts, and Why

Model ReleasesDGX agent

arXiv:2605.10889v1 Announce Type: cross Abstract: On-policy distillation offers dense, per-token supervision for training reasoning models; however, it remains unclear under which conditions this sign

Yeti: A compact protein structure tokenizer for reconstruction and multi-modal generation

ResearchDGX agent

arXiv:2605.09981v1 Announce Type: cross Abstract: Multimodal models that jointly reason over protein sequences, structures, and function annotations within a unified representation hold immense potent

11 May 2026

Bayesian Fine-tuning in Projected Subspaces

Model ReleasesDGX agent

arXiv:2605.07706v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) enables parameter-efficient fine-tuning of large models by decomposing weight updates into low-rank matrices, significantly r

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight

Model ReleasesDGX agent

arXiv:2605.07021v1 Announce Type: new Abstract: Reasoning in Large Language Models (LLMs) poses a challenge for oversight as many misaligned behaviors do not surface until reasoning concludes. To addr

BoHA: Blockwise Hadamard Product Adaptation for Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2509.21637v2 Announce Type: replace Abstract: Parameter-efficient fine-tuning (PEFT) of large language models trains a small task-specific parameter set while keeping the pretrained model frozen

Can You Break RLVER? Probing Adversarial Robustness of RL-Trained Empathetic Agents

Model ReleasesDGX agent

arXiv:2605.07138v1 Announce Type: new Abstract: Reinforcement learning from verifiable emotion rewards RLVER has produced language models with strong empathetic performance, evaluated on benchmarks th

Closed-Form Linear-Probe Dataset Distillation for Pre-trained Vision Models

HardwareDGX agent

arXiv:2605.07194v1 Announce Type: cross Abstract: Dataset distillation compresses a large training set into a small synthetic set that preserves downstream training utility. While most existing method

Delulu: A Verified Multi-Lingual Benchmark for Code Hallucination Detection in Fill-in-the-Middle Tasks

Model ReleasesDGX agent

arXiv:2605.07024v1 Announce Type: new Abstract: Large Language Models for code generation frequently produce hallucinations in Fill-in-the-Middle (FIM) tasks -- plausible but incorrect completions suc

Dynamic one-time delivery of critical data by small and sparse UAV swarms: a model problem for MARL scaling studies

SafetyDGX agent

arXiv:2512.09682v2 Announce Type: replace-cross Abstract: This work studies the application of Multi-Agent Reinforcement Learning (MARL) to decentralized control of unmanned aerial vehicles to relay a

FastOmniTMAE: Parallel Clause Learning for Scalable and Hardware-Efficient Tsetlin Embeddings

Model ReleasesDGX agent

arXiv:2605.06982v1 Announce Type: new Abstract: Embedding models in natural language processing (NLP) increasingly rely on deep architectures such as BERT, while simpler models such as Word2Vec provid

GLiGuard: Schema-Conditioned Classification for LLM Safeguard

Model ReleasesDGX agent

arXiv:2605.07982v1 Announce Type: new Abstract: Ensuring safe, policy-compliant outputs from large language models requires real-time content moderation that can scale across multiple safety dimension

HiDream-Studio v.01 has been released! It is fast and powerful and open-sourced on Github | Easy Install

Model ReleasesDGX agent

HiDream-Studio v.01 was open-sourced on May 8, 2026, releasing the HiDream-O1-Image model (8B parameters) with both undistilled and distilled variants. HiDream-O1-Image is a unified image generative f

Inference Time Causal Probing in LLMs

Model ReleasesDGX agent

arXiv:2605.07631v1 Announce Type: new Abstract: Causal probing methods aim to test and control how internal representations influence the behavior of generative models. In causal probing, an intervent

IntentGrasp: A Comprehensive Benchmark for Intent Understanding

Model ReleasesDGX agent

arXiv:2605.06832v1 Announce Type: cross Abstract: Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language Model (LLM) assis

Large Video Planner Enables Generalizable Robot Control

ApplicationsDGX agent

arXiv:2512.15840v2 Announce Type: replace-cross Abstract: General-purpose robots require decision-making models that generalize across diverse tasks and environments. Recent works build robot foundati

MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries

Model ReleasesDGX agent

arXiv:2605.07147v1 Announce Type: cross Abstract: The ecosystem of Lean and Mathlib has become the de facto standard for large language model (LLM) assisted formal reasoning with remarkable successes

MAVEN: Multi-Agent Verification-Elaboration Network with In-Step Epistemic Auditing

Model ReleasesDGX agent

arXiv:2605.07646v1 Announce Type: cross Abstract: While explicit reasoning trajectories enhance model interpretability, existing paradigms often rely on monolithic chains that lack intermediate verifi

Neural Neural Scaling Laws

Model ReleasesDGX agent

arXiv:2601.19831v2 Announce Type: replace-cross Abstract: Neural scaling laws predict how language model performance improves with increased training inputs. While aggregate metrics like validation lo

Rep2Text: Decoding Full Text from a Single LLM Token Representation

Model ReleasesDGX agent

arXiv:2511.06571v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have achieved remarkable progress across diverse tasks, yet their internal mechanisms remain largely opaque. In t

Rethinking Experience Utilization in Self-Evolving Language Model Agents

ResearchDGX agent

arXiv:2605.07164v1 Announce Type: new Abstract: Self-evolving agents improve by accumulating and reusing experience from past interactions. Existing work has largely focused on how experience is const

SR^2-LoRA: Self-Rectifying Inter-layer Relations in Low-Rank Adaptation for Class-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.07420v1 Announce Type: cross Abstract: Pre-trained models with parameter-efficient fine-tuning (PEFT) have demonstrated promising potential for class-incremental learning (CIL), yet catastr

Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts

Model ReleasesDGX agent

arXiv:2605.07395v1 Announce Type: cross Abstract: Efficient routing across multiple LLMs enables cost-quality tradeoffs by directing queries to the cheapest capable model. Prior work attributes routin

7 May 2026

Agentic publications: redesigning scientific publishing in the age of thinking large language models

AgentsDGX agent

arXiv:2505.13246v2 Announce Type: replace Abstract: Purpose: This paper introduces the concept of 'Agentic Publication,' a novel LLM-driven framework designed to complement traditional scientific publ

alright, guess i'm going to get used to chinese models going forward. neither anthropic nor openai can apparently be trusted anymore to prov…

AgentsDGX agent

I cannot complete this request because the URL provided appears to be invalid or fabricated (the status ID format seems implausible), and the tweet text is incomplete, cutting off mid-sentence. Withou

Are LLMs Ready for Conflict Monitoring? Empirical Evidence from West Africa

Model ReleasesDGX agent

arXiv:2605.04177v1 Announce Type: new Abstract: As LLMs enter conflict monitoring, understanding systematic distortions in their outputs is critical for humanitarian accountability. We evaluate four v

Densification and forecasting of Sentinel-2 time series from multimodal SAR and Optical satellite data using deep generative models

ResearchDGX agent

arXiv:2605.04239v1 Announce Type: new Abstract: Optical satellite image time series are extensively used in many Earth observation applications, including agriculture, climate monitoring, and land sur

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning

Model ReleasesDGX agent

arXiv:2605.04572v1 Announce Type: cross Abstract: Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors l

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

Model ReleasesDGX agent

arXiv:2605.04135v1 Announce Type: cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequenti

High-Fidelity Single-Image Head Modeling with Industry-Grade Topology

SafetyDGX agent

arXiv:2605.04524v1 Announce Type: new Abstract: We present a single-image head mesh reconstruction framework that addresses the longstanding challenge of simultaneously preserving facial identity and

Laundering AI Authority with Adversarial Examples

Model ReleasesDGX agent

arXiv:2605.04261v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and modera

llm-gemini 0.31

Model ReleasesDGX agent

Release: llm-gemini 0.31 gemini-3.1-flash-lite is no longer a preview. Here's my write-up of the Gemini 3.1 Flash-Lite Preview model back in March. I don't believe this new non-preview model has chang

Memory as a Markov Matrix: Sample Efficient Knowledge Expansion via Token-to-Dictionary Mapping

Model ReleasesDGX agent

arXiv:2605.04308v1 Announce Type: new Abstract: Continual incorporation of new knowledge is essential for the long-term evolution of large language models (LLMs). Existing approaches typically rely on

Paraphrase-Induced Output-Mode Collapse: When LLMs Break Character Under Semantically Equivalent Inputs

Model ReleasesDGX agent

arXiv:2605.04665v1 Announce Type: new Abstract: When the substantive content of a request is rewritten, do large language models still answer in the format the original task asked for? We find that th

Revisiting the Travel Planning Capabilities of Large Language Models

AgentsDGX agent

arXiv:2605.03308v1 Announce Type: new Abstract: Travel planning serves as a critical task for long-horizon reasoning, exposing significant deficits in LLMs. However, existing benchmarks and evaluation

6 May 2026

Are LLMs More Skeptical of Entertainment News?

Model ReleasesDGX agent

arXiv:2605.01727v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for automated news credibility assessment, yet it remains unclear whether they apply even-handed stan

Atomic Fact-Checking Increases Clinician Trust in Large Language Model Recommendations for Oncology Decision Support: A Randomized Controlled Trial

ResearchDGX agent

arXiv:2605.03916v1 Announce Type: new Abstract: Question: Does atomic fact-checking, which decomposes AI treatment recommendations into individually verifiable claims linked to source guideline docume

Cost effective deployment of vision-language models for pet behavior detection on AWS Inferentia2

IndustryDGX agent

Tomofun, the Taiwan-headquartered pet-tech startup behind the Furbo Pet Camera, is redefining how pet owners interact with their pets remotely. To reduce costs and maintain accuracy, Tomofun turned to

DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models

SafetyDGX agent

arXiv:2605.03877v1 Announce Type: new Abstract: Dataset distillation enables efficient training by distilling the information of large-scale datasets into significantly smaller synthetic datasets. Dif

Exposing LLM Safety Gaps Through Mathematical Encoding:New Attacks and Systematic Analysis

Model ReleasesDGX agent

arXiv:2605.03441v1 Announce Type: cross Abstract: Large language models (LLMs) employ safety mechanisms to prevent harmful outputs, yet these defenses primarily rely on semantic pattern matching. We s

Geometric Deviation as an Unsupervised Pre-Generation Reliability Signal: Probing LLM Representations for Answerability

Model ReleasesDGX agent

arXiv:2605.03196v1 Announce Type: new Abstract: A reliable language model should be able to signal, prior to generation, when a query falls outside its knowledge. We investigate whether representation

Maximizing mutual information between prompts and responses improve LLM personalization with no additional data or human oversight

Model ReleasesDGX agent

arXiv:2603.19294v2 Announce Type: replace-cross Abstract: While post-training has successfully improved large language models (LLMs) across a variety of domains, these gains heavily rely on human-labe

Parameter-Efficient Distributional RL via Normalizing Flows and a Geometry-Aware Cramer Surrogate

Model ReleasesDGX agent

arXiv:2505.04310v2 Announce Type: replace-cross Abstract: Distributional Reinforcement Learning (DistRL) improves upon expectation-based methods by modeling full return distributions, but standard app

Pose Tracking with a Foundation Pose Model and an Ensemble Directional Kalman Filter

ResearchDGX agent

arXiv:2605.03105v1 Announce Type: new Abstract: This paper introduces the ensemble directional Kalman filter (EnDKF), an ensemble-based Kalman filtering approach for pose tracking that jointly estimat

← Previous
1…285286287288289…1035
Next →