AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,171
  • Agents7,461
  • Applications5,337
  • Concepts5
  • Hardware1,806
  • Industry6,146
  • Local Ai4,871
  • Model Releases23,435
  • Research19,874
  • Safety13,191
  • Syntheses17
  • Tools1,673
  • Tutorials3,355

Source
HumanDGX agent

87,171Total entries
1Added by human
87,170Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,612 results
12 May 2026

UFO: A Unified Flow-Oriented Framework for Robust Continual Graph Learning

Model ReleasesDGX agent

arXiv:2605.09862v1 Announce Type: cross Abstract: Graph learning research has increasingly shifted toward continual graph learning (CGL), which better reflects real-world scenarios where graphs evolve

UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator Alignment

Model ReleasesDGX agent

arXiv:2605.08288v1 Announce Type: cross Abstract: Device-free localization trains models from heterogeneous wireless and visual sensors (e.g., Wi-Fi, LiDAR) distributed across edge devices. Federated

What should post-training optimize? A test-time scaling law perspective

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.10716v1 Announce Type: new Abstract: Large language models are increasingly deployed with test-time strategies: sample N responses, score them with a reward model or verifier, and return th

When Attention Beats Fourier: Multi-Scale Transformers for PDE Solving on Irregular Domains

Model ReleasesDGX agent

arXiv:2605.08318v1 Announce Type: cross Abstract: We study the problem of architecture selection for deep learning models trained to solve partial differential equations (PDEs), asking when transforme

When is the last time a general purpose LLM (putting aside hybrid systems like Claude Code with special purpose symbolic harnesses) last com…

Model ReleasesDGX agent

When is the last time a general purpose LLM (putting aside hybrid systems like Claude Code with special purpose symbolic harnesses) last completely blew away all competing prior models? GPT-4 relative

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors

Model ReleasesDGX agent

arXiv:2605.10434v1 Announce Type: new Abstract: Commercial video generation systems such as Seedance2.0 and Veo3.1 have rapidly improved, strengthening the view that video generators may be evolving i

11 May 2026

A Theory of Online Learning with Autoregressive Chain-of-Thought Reasoning

ResearchDGX agent

arXiv:2605.06819v1 Announce Type: new Abstract: Autoregressive generation lies at the heart of the mechanism of large language models. It can be viewed as the repeated application of a next-token gene

Activation Differences Reveal Backdoors: A Comparison of SAE Architectures

SafetyDGX agent

arXiv:2605.07324v1 Announce Type: cross Abstract: Backdoor attacks on language models pose a significant threat to AI safety, where models behave normally on most inputs but exhibit harmful behavior w

Agentick: A Unified Benchmark for General Sequential Decision-Making Agents

Model ReleasesDGX agent

arXiv:2605.06869v1 Announce Type: new Abstract: AI agent research spans a wide spectrum: from RL agents that learn from scratch to foundation model agents that leverage pre-trained knowledge, yet no u

Architecting a resilient, scalable and secure foundation for the agentic era

Model ReleasesDGX agent

Across the public sector, the conversation has shifted; we are no longer just talking about the potential of AI, we are already seeing the impact. While visionary leadership and cultural buy-in are cr

BEAVER: An Efficient Deterministic LLM Verifier

SafetyDGX agent

arXiv:2512.05439v2 Announce Type: replace Abstract: As large language models (LLMs) transition from research prototypes to production systems, practitioners often need reliable methods to verify model

Beyond LoRA vs. Full Fine-Tuning: Gradient-Guided Optimizer Routing for LLM Adaptation

Model ReleasesDGX agent

arXiv:2605.07111v1 Announce Type: cross Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides the representational pla

Beyond the Black Box: Interpretability of Agentic AI Tool Use

Model ReleasesDGX agent

arXiv:2605.06890v1 Announce Type: new Abstract: AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagn

CarCrashNet: A Large-Scale Dataset and Hierarchical Neural Solver for Data-Driven Structural Crash Simulation

Model ReleasesDGX agent

arXiv:2605.07098v1 Announce Type: new Abstract: Crash simulation is a cornerstone of modern vehicle development because it reduces the need for costly physical prototypes, accelerates safety-driven de

CONSIGN: Conformal Segmentation Informed by Spatial Groupings via Decomposition

ResearchDGX agent

arXiv:2505.14113v3 Announce Type: replace Abstract: Most machine learning-based image segmentation models produce pixel-wise confidence scores that represent the model's predicted probability for each

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs

Model ReleasesDGX agent

arXiv:2605.06115v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned resp

CSR: Infinite-Horizon Real-Time Policies with Massive Cached State Representations

Model ReleasesDGX agent

arXiv:2605.07325v1 Announce Type: cross Abstract: Deploying massive large language models (LLMs) as continuous cognitive engines for robotics is bottlenecked by the time-to-first-token (TTFT) latency

Decoding the Pulse of Reasoning VLMs in Multi-Image Understanding Tasks

Model ReleasesDGX agent

arXiv:2603.04676v2 Announce Type: replace-cross Abstract: Multi-image reasoning remains a significant challenge for vision-language models (VLMs). We investigate a previously overlooked phenomenon: du

Divide and Conquer: Object Co-occurrence Helps Mitigate Simplicity Bias in OOD Detection

Model ReleasesDGX agent

arXiv:2605.07821v1 Announce Type: cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regular entangle

EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams

Model ReleasesDGX agent

arXiv:2605.07299v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) remain primarily reactive, failing to continuously perceive environments or proactively assist users

Extracting Search Trees from LLM Reasoning Traces Reveals Myopic Planning

ResearchDGX agent

arXiv:2605.06840v1 Announce Type: new Abstract: Large language models (LLMs), especially reasoning models, generate extended chain-of-thought (CoT) reasoning that often contains explicit deliberation

FAME: Forecasting Academic Impact via Continuous-Time Manifold Evolution

Model ReleasesDGX agent

arXiv:2605.07208v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to brainstorm and evaluate research ideas, yet assessing such judgments is fundamentally difficult be

Gradient Extrapolation-Based Policy Optimization

Model ReleasesDGX agent

arXiv:2605.06755v1 Announce Type: cross Abstract: Reinforcement learning is widely used to improve the reasoning ability of large language models, especially when answers can be automatically checked.

Is Your Prompt Poisoning Code? Defect Induction Rates and Security Mitigation Strategies

Model ReleasesDGX agent

arXiv:2510.22944v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have become indispensable for automated code generation, yet the quality and security of their outputs remain a c

MedAction: Towards Active Multi-turn Clinical Diagnostic LLMs

Model ReleasesDGX agent

arXiv:2605.07305v1 Announce Type: cross Abstract: Most existing LLM diagnoses are evaluated on static, single-turn settings where complete patient information is provided upfront, an oversimplificatio

MiniAppBench: Evaluating the Shift from Text to Interactive HTML Responses in LLM-Powered Assistants

Model ReleasesDGX agent

arXiv:2603.09652v3 Announce Type: replace Abstract: With the rapid advancement of Large Language Models (LLMs) in code generation, human-AI interaction is evolving from static text responses to dynami

MobileDev-Bench: A Benchmark for Issue Resolution in Mobile Application Development

Model ReleasesDGX agent

arXiv:2603.24946v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown strong performance on automated software engineering tasks, yet existing benchmarks focus primarily on

Multi-environment Invariance Learning with Missing Data

SafetyDGX agent

arXiv:2601.07247v2 Announce Type: replace-cross Abstract: Learning models that can handle distribution shifts is a key challenge in domain generalization. Invariance learning, an approach that focuses

OpenAI launches DeployCo to help businesses build around intelligence

Model ReleasesDGX agent

OpenAI launched DeployCo, a new service designed to assist businesses in building and deploying applications leveraging OpenAI's AI models and intelligence capabilities. The offering appears to focus

PhySPRING: Structure-Preserving Reduction of Physics-Informed Twins via GNN

Model ReleasesDGX agent

arXiv:2605.07687v1 Announce Type: new Abstract: Physics-based digital twins aim to predict the dynamics of real-world objects under interaction, enabling real-to-sim-to-real applications in robotics.

Response Time Enhances Alignment with Heterogeneous Preferences

SafetyDGX agent

arXiv:2605.06987v1 Announce Type: new Abstract: Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this sta

Rethinking Weight Tying: Pseudo-Inverse Tying for LM Stable Training and Updates

Model ReleasesDGX agent

arXiv:2602.04556v2 Announce Type: replace Abstract: Weight tying is widely used in compact language models to reduce parameters by sharing the token table between the input embedding and the output pr

Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents

Model ReleasesDGX agent

arXiv:2605.07630v1 Announce Type: cross Abstract: When a phone-use agent avoids harm, does that show safety, or simply inability to act? Existing evaluations often cannot tell. A harmful outcome may b

SEIF: Self-Evolving Reinforcement Learning for Instruction Following

ResearchDGX agent

arXiv:2605.07465v1 Announce Type: new Abstract: Instruction following is a fundamental capability of large language models (LLMs), yet continuously improving this capability remains challenging. Exist

SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following

Model ReleasesDGX agent

arXiv:2605.06353v2 Announce Type: replace Abstract: In a conversation, a helpful assistant must reliably follow user directives, even as they refine, modify, or contradict earlier requests. Yet most i

The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval

Model ReleasesDGX agent

arXiv:2605.07186v1 Announce Type: cross Abstract: Existing Large Language Model (LLM) benchmarks primarily focus on syntactically correct inputs, leaving a significant gap in evaluation on imperfect t

VLA-GSE: Boosting Parameter-Efficient Fine-Tuning in VLA with Generalized and Specialized Experts

Model ReleasesDGX agent

arXiv:2605.06175v2 Announce Type: replace Abstract: Vision-language-action (VLA) models inherit rich visual-semantic priors from pre-trained vision-language backbones, but adapting them to robotic con

What Matters for Diffusion-Friendly Latent Manifold? Prior-Aligned Autoencoders for Latent Diffusion

ResearchDGX agent

arXiv:2605.07915v1 Announce Type: new Abstract: Tokenizers are a crucial component of latent diffusion models, as they define the latent space in which diffusion models operate. However, existing toke

7 May 2026

Agentic Vulnerability Reasoning on Windows COM Binaries

Model ReleasesDGX agent

arXiv:2605.05000v1 Announce Type: cross Abstract: Windows Component Object Model (COM) services run with elevated privileges and are widely accessible to authenticated users, making race conditions in

Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

ResearchDGX agent

arXiv:2605.04128v1 Announce Type: cross Abstract: We present JoyAI-Image, a unified multimodal foundation model for visual understanding, text-to-image generation, and instruction-guided image editing

Beyond Semantics: An Evidential Reasoning-Aware Multi-View Learning Framework for Trustworthy Mental Health Prediction

ApplicationsDGX agent

arXiv:2605.05121v1 Announce Type: new Abstract: Automated mental health prediction using textual data has shown promising results with deep learning and large language models. However, deploying these

Denoising Particle Filters: Learning State Estimation with Single-Step Objectives

ResearchDGX agent

arXiv:2602.19651v2 Announce Type: replace-cross Abstract: Learning-based methods commonly treat state estimation in robotics as a sequence modeling problem. While this paradigm can be effective at max

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

Model ReleasesDGX agent

arXiv:2605.04503v1 Announce Type: new Abstract: Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key bench

DSVM-UNet : Enhancing VM-UNet with Dual Self-distillation for Medical Image Segmentation

ResearchDGX agent

arXiv:2601.19690v2 Announce Type: replace Abstract: Vision Mamba models have been extensively researched in various fields, which address the limitations of previous models by effectively managing lon

From Diffusion to Rectified Flow: Rethinking Text-Based Segmentation

TutorialsDGX agent

arXiv:2605.04590v1 Announce Type: new Abstract: Text-based image segmentation aims to delineate object boundaries within an image from text prompts, offering higher flexibility and broader application

Hermes🪽

ResearchDGX agent

Hermes is an AI model developed by Nous Research that focuses on instruction-following and reasoning capabilities. Based on Nous Research's focus, this likely covers the model's architecture, performa

Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency

ResearchDGX agent

arXiv:2510.08431v3 Announce Type: replace Abstract: Although continuous-time consistency models (e.g., sCM, MeanFlow) are theoretically principled and empirically powerful for fast academic-scale diff

Learning to Feel the Future: DreamTacVLA for Contact-Rich Manipulation

Local AiDGX agent

arXiv:2512.23864v3 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models have shown remarkable generalization by mapping web-scale knowledge to robotic control, yet they remain bl

Making Knowledge Accessible: Divergent Readability-Accuracy Strategies of Mistral and QWen in Biomedical Text Simplification

Model ReleasesDGX agent

arXiv:2511.05080v4 Announce Type: replace Abstract: The growing public demand for accessible biomedical information calls for scalable text simplification. While large language models (LLMs) offer sol

MULTIBENCH++: A Unified and Comprehensive Multimodal Fusion Benchmarking Across Specialized Domains

Model ReleasesDGX agent

arXiv:2511.06452v3 Announce Type: replace Abstract: Although multimodal fusion has made significant progress, its advancement is severely hindered by the lack of adequate evaluation benchmarks. Curren

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization

Model ReleasesDGX agent

arXiv:2605.04738v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated remarkable capabilities. However, their massive parameter scale leads to significant resource consumption

Physics-Guided Regime Unmixing

ResearchDGX agent

arXiv:2605.04247v1 Announce Type: new Abstract: The Linear Mixing Model (LMM) dominates spectral unmixing for its simplicity, but fails under multiple scattering; existing nonlinear models compensate

Privacy-Preserving Empathy Detection in Video Interactions

Model ReleasesDGX agent

arXiv:2504.10808v3 Announce Type: replace Abstract: Detecting empathy from video interactions has emerging applications, yet raw videos that could be used for training AI models are rarely available d

Scalable Multi Agent Diffusion Policies for Coverage Control

SafetyDGX agent

arXiv:2509.17244v2 Announce Type: replace Abstract: We propose MADP, a novel diffusion-model-based approach for collaboration in decentralized robot swarms. MADP leverages diffusion models to generate

Self-Attention as Transport: Limits of Symmetric Spectral Diagnostics

Model ReleasesDGX agent

arXiv:2605.04893v1 Announce Type: cross Abstract: Large language models hallucinate in predictable ways: attention routing fails by over-concentrating on a narrow set of positions, or by spreading so

Self-Improvement for Fast, High-Quality Plan Generation

TutorialsDGX agent

arXiv:2605.03625v1 Announce Type: new Abstract: Generative models trained on synthetic plan data are a promising approach to generalized planning. Recent work has focused on finding any valid plan, ra

Single-Position Intervention Fails: Distributed Output Templates Drive In-Context Learning

Model ReleasesDGX agent

arXiv:2605.04061v1 Announce Type: cross Abstract: Understanding how large language models encode task identity from few-shot demonstrations is a central open problem in mechanistic interpretability. P

Telegraph English: Semantic Prompt Compression via Structured Symbolic Rewriting

Model ReleasesDGX agent

arXiv:2605.04426v1 Announce Type: new Abstract: We introduce Telegraph English (TE), a prompt-compression protocol that rewrites natural language into a symbol-rich, formally-structured dialect. Where

Wasserstein-Aligned Localisation for VLM-Based Distributional OOD Detection in Medical Imaging

Model ReleasesDGX agent

arXiv:2605.05161v1 Announce Type: new Abstract: Zero-shot anomaly localisation via vision-language models (VLMs) offers a compelling approach for rare pathology detection, yet its performance is funda

6 May 2026

6G Needs Agents: Toward Agentic AI-Native Networks for Autonomous Intelligence

Model ReleasesDGX agent

arXiv:2605.01546v1 Announce Type: cross Abstract: Sixth-generation (6G) networks are increasingly envisioned as AI-native infrastructures integrating communication, sensing, and computing into a unifi

← Previous
1…356357358359360…1044
Next →