AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,814
  • Agents7,519
  • Applications5,378
  • Concepts5
  • Hardware1,822
  • Industry6,162
  • Local Ai4,908
  • Model Releases23,658
  • Research20,008
  • Safety13,291
  • Syntheses17
  • Tools1,674
  • Tutorials3,372

Source
HumanDGX agent

87,814Total entries
1Added by human
87,813Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,138 results
3 Jun 2026

Learning Self-Interpretation from Interpretability Artifacts: Training Lightweight Adapters on Vector-Label Pairs

SafetyDGX agent

arXiv:2602.10352v2 Announce Type: replace-cross Abstract: Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitiv

MemoGen: Can Past Experience Improve Future Text-to-Image Generation?

Model ReleasesDGX agent

arXiv:2606.03243v1 Announce Type: new Abstract: Modern text-to-image models have achieved strong visual synthesis, yet remain unreliable when prompts require implicit visual constraints, relational re

Message Tuning Outshines Graph Prompt Tuning: A Prismatic Space Perspective

Model ReleasesDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.03290v1 Announce Type: cross Abstract: Graph Foundation Models (GFMs), built upon the Pre-training and Adaptation paradigm, have emerged as a research hotspot in graph learning. For GNN-bas

MiMo v2.5 (pro) availability

Local AiDGX agent

MiMo-V2.5-Pro is a model available on Hugging Face that was requested to be added to Ollama's cloud models in May 2026. The discussion on r/ollama likely covers the availability status of this Xiaomi-

Multi^2: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments

Model ReleasesDGX agent

arXiv:2606.03698v1 Announce Type: new Abstract: A central goal of large language model (LLM) research is to build agentic systems that can plan, act, and adapt through sustained interaction with dynam

Nanocoder 1.27.0 - skills, daemon + more 🔥

Local AiDGX agent

Nanocoder 1.27.0 is an agentic coding tool available in your terminal that runs on any AI model you choose, whether local models via Ollama or cloud providers like OpenAI and Anthropic. This release i

NeuroArmor: Safe-Variant-Guided Representation Consistency for Selective Re-Anchoring in Jailbreak Defense

Model ReleasesDGX agent

arXiv:2606.03486v1 Announce Type: cross Abstract: Large language models remain vulnerable to jailbreak attacks that hide harmful intent behind seemingly ordinary requests such as role-play, translatio

OpenEAI-Platform: An Open-source Embodied Artificial Intelligence Hardware-Software Unified Platform

Model ReleasesDGX agent

arXiv:2606.03392v1 Announce Type: new Abstract: Embodied AI in the real world requires both accurate hardware and robust vision-language-action (VLA) policies. We present OpenEAI-Platform, a fully ope

Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

Model ReleasesDGX agent

arXiv:2606.03236v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have substantially advanced mobile agents, yet proactive mobile assistance remains challenging because agents m

PrimeSVT: An Automated Memory-aware Pruning Framework with Prioritized Compression Policy for Spiking Vision Transformers

SafetyDGX agent

arXiv:2606.03428v1 Announce Type: cross Abstract: The large sizes of Spiking Vision Transformers (SViTs) still hinder their embedded implementation, highlighting the need for model compression. State-

Proof-Refactor: Refactoring Generated Formal Proofs into Modular Artifacts

Model ReleasesDGX agent

arXiv:2606.03743v1 Announce Type: new Abstract: While Large Language Models (LLMs) have shown strong performance in generating formal proofs, their outputs often remain less readable, modular, maintai

Qwen-Image-Flash: Beyond Objective Design

Model ReleasesDGX agent

arXiv:2606.03746v1 Announce Type: cross Abstract: Few-step distillation has become an effective strategy for accelerating advanced visual generative models, yet prior work has largely focused on disti

Samudra 2: Scaling Ocean Emulators across Resolutions

Model ReleasesDGX agent

arXiv:2606.02610v1 Announce Type: cross Abstract: Ocean general circulation models (OGCMs) are essential to climate science but computationally expensive, limiting ensemble size and forcing scenarios.

Scalable On-Hardware Training of Quantum Neural Networks and Application to Clinical Data Imputation

Model ReleasesDGX agent

arXiv:2606.03517v1 Announce Type: cross Abstract: Training quantum neural networks (QNNs) on quantum hardware is currently bottlenecked by the cost of gradient estimation: standard parameter-shift met

SEA-NLI: Natural Language Inference as a Lens into Southeast Asian Cultural Understanding

Model ReleasesDGX agent

arXiv:2606.03284v1 Announce Type: new Abstract: Frontier LLMs perform well in Western contexts, but remain poorly tested on underrepresented cultures such as those in Southeast Asia (SEA). Existing NL

SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos

Model ReleasesDGX agent

arXiv:2606.02745v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-speci

SynCred-Bench: Benchmarking Synthetic Credibility in AI-Generated Visual Misinformation

Model ReleasesDGX agent

arXiv:2606.03348v1 Announce Type: cross Abstract: Recent generative models can now produce visual artifacts with realistic embedded text and layouts, creating a new misinformation threat: synthetic cr

Trading Human Curation for Synthetic Augmentation in RLVR

Model ReleasesDGX agent

arXiv:2606.03800v1 Announce Type: cross Abstract: The supply of high-quality training tasks is a central bottleneck for reinforcement learning from verifiable rewards (RLVR) on agentic language models

TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment

Model ReleasesDGX agent

arXiv:2606.03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-w

VidMsg: A Benchmark for Implicit Message Inference in Short Videos

Model ReleasesDGX agent

arXiv:2606.03635v1 Announce Type: cross Abstract: Understanding short online videos involves more than identifying visible objects and actions; video makers often include an underlying message or purp

Weak Diffusion Priors Can Still Achieve Strong Inverse-Problem Performance

ResearchDGX agent

arXiv:2601.22443v2 Announce Type: replace-cross Abstract: Can a diffusion model trained on bedrooms recover human faces? Diffusion models are widely used as priors for inverse problems, but standard a

What Do Students Learn? A Feature-Level Analysis of Dark Knowledge

TutorialsDGX agent

arXiv:2606.03052v1 Announce Type: new Abstract: Knowledge Distillation (KD) is a powerful tool for model compression, yet the precise mechanisms by which student models acquire feature representations

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

Model ReleasesDGX agent

arXiv:2606.02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-regis

X-RAY: Mapping LLM Reasoning Capability via Formalized and Calibrated Probes

ResearchDGX agent

arXiv:2603.05290v2 Announce Type: replace Abstract: Large language models (LLMs) achieve promising performance, yet their ability to reason remains poorly understood. Existing evaluations largely emph

2 Jun 2026

A Closer Look at In-Distribution vs. Out-of-Distribution Accuracy for Open-Set Test-time Adaptation

Model ReleasesDGX agent

arXiv:2606.01973v1 Announce Type: cross Abstract: Open-set test-time adaptation (TTA) updates models on new data in the presence of input shifts and unknown output classes. While recent methods have m

A Machine-to-Machine Knowledge-Guided LLM Agent for Generalizable Radiotherapy Treatment Planning

Model ReleasesDGX agent

arXiv:2606.00922v1 Announce Type: cross Abstract: In this work, we propose a prototype machine-to-machine (M2M) knowledge-guided Large Language Model (LLM) framework for automated radiotherapy treatme

A unifying Bayesian framework for adversarial robustness

ResearchDGX agent

arXiv:2510.09288v2 Announce Type: replace-cross Abstract: The vulnerability of machine learning models to adversarial attacks remains a critical societal security challenge. Traditional defenses, such

Anthropic expands Project Glasswing cybersecurity program to 150 more organizations

Model ReleasesDGX agent

Anthropic PBC is expanding a program that enables organizations to test their cybersecurity defenses using its Claude Mythos Preview model. The initiative, which is known as Project Glasswing, launche

ASKD-Whisper: Adaptive Self-knowledge Distillation for Efficient and Low-Latency Automatic Speech Recognition

ResearchDGX agent

arXiv:2601.19919v2 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is one of the most effective paradigms for compressing large-scale foundation models into deployable architectures

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

Model ReleasesDGX agent

arXiv:2606.02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural

AXIOM: A Trust-First Neuro-Symbolic Execution Architecture for Verifiable Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2606.00671v1 Announce Type: new Abstract: We present AXIOM, a trust-first neuro-symbolic execution architecture for natural-language mathematical reasoning. In AXIOM, the language model function

Better with Experience: Self-Evolving LLM Agents for Evidence-Grounded Health Community Notes

Model ReleasesDGX agent

arXiv:2606.02215v1 Announce Type: new Abstract: Large Language Model (LLM)-augmented Community Notes offer a scalable path for timely, evidence-grounded correction of health misinformation on social p

Beyond Rigid: Benchmarking Non-Rigid Video Editing

Model ReleasesDGX agent

arXiv:2601.18340v2 Announce Type: replace Abstract: As video generation models are increasingly expected to manipulate physical dynamics, there is a growing need to move evaluation beyond appearance f

Beyond Scalar Rewards: Dense Feedback for LLM Policy Synthesis in Sequential Social Dilemmas

Model ReleasesDGX agent

arXiv:2603.19453v2 Announce Type: replace Abstract: We study LLM policy synthesis: using a language model to iteratively generate programmatic agent policies for multi-agent environments. Rather than

CART: Context-Anchored Recurrent Transformer -- A Parameter-Efficient Architecture with Learned Stability

Model ReleasesDGX agent

arXiv:2606.01495v1 Announce Type: cross Abstract: We present CART (Context-Anchored Recurrent Transformer), a parameter-efficient language model that reuses a single shared core block R times across d

Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs

Model ReleasesDGX agent

arXiv:2603.24511v2 Announce Type: replace-cross Abstract: We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on whi

Codex is becoming a productivity tool for everyone

Model ReleasesDGX agent

OpenAI's Codex is evolving beyond code generation to become a general productivity tool accessible to non-programmers for knowledge work tasks. The tool leverages large language models to assist with

Continuous Reasoning for Vision-Language-Action

ResearchDGX agent

arXiv:2606.00229v1 Announce Type: cross Abstract: Natural language is a powerful reasoning medium for language and vision-language models, but it is mismatched to the granularity of continuous control

ContinuousBench: Can Differentially Private Synthetic Text Improve Capabilities?

Model ReleasesDGX agent

arXiv:2606.01849v1 Announce Type: cross Abstract: Differentially private (DP) text synthesis promises to unlock sensitive corpora for model training, but it remains unclear whether DP synthetic data t

CRAM: Centroid-Routing and Adaptive MoE for Multimodal Continual Instruction Tuning

Model ReleasesDGX agent

arXiv:2606.02502v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) unify heterogeneous vision-language tasks under a shared generative framework via instruction tuning, yet real-

CRMA: A Spectrally-Bounded Backbone for Modular Continual Fine-Tuning of LLMs

Model ReleasesDGX agent

arXiv:2606.00382v1 Announce Type: new Abstract: Sequential fine-tuning of large language models forces a choice: let the shared substrate keep learning and accept catastrophic forgetting, or freeze it

Cross-Generational Transfer of Adversarial Attacks Reveals Non-Monotonic Safety Alignment in LLMs

Model ReleasesDGX agent

arXiv:2606.00813v1 Announce Type: cross Abstract: Safety alignment in LLMs does not improve monotonically across model generations. Studying four generations of Google's Gemma family (7B-31B) with qua

CV-Arena: An Open Benchmark for Instructional Computer Vision Problem Solving with Human-AI Collaborative Preferences

Model ReleasesDGX agent

arXiv:2606.00931v1 Announce Type: cross Abstract: Instruction-guided image editing is becoming a general interface for visual work, yet existing benchmarks still focus largely on narrow appearance edi

Data Collection for Training Quality-Control AI in Carpet Manufacturing

Model ReleasesDGX agent

arXiv:2606.01023v1 Announce Type: cross Abstract: Visual inspection remains the dominant quality-control practice in woven and tufted carpet production, yet it is slow, subjective, and inconsistent at

DenseMLLM: Standard Multimodal LLMs for Dense Prediction

ResearchDGX agent

arXiv:2602.14134v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated exceptional capabilities in high-level visual understanding. However, extending the

Dialectics of Alignment: Harnessing Unsafe Knowledge for Dynamic Safety Routing

SafetyDGX agent

arXiv:2606.00686v1 Announce Type: new Abstract: The prevailing paradigm in large language model (LLM) alignment operates via erasure, filtering unsafe data or training models to strictly refuse harmfu

DOT-MoE: Differentiable Optimal Transport for MoEfication

SafetyDGX agent

arXiv:2606.01666v1 Announce Type: cross Abstract: The scaling of Large Language Models (LLMs) has driven significant performance gains but created substantial challenges in inference efficiency. While

Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing

Model ReleasesDGX agent

arXiv:2606.01393v1 Announce Type: cross Abstract: Document parsing and recognition are fundamental capabilities for vision-language models (VLMs) and document processing systems. However, existing Opt

Echo State Networks for Time Series Forecasting: Hyperparameter Sweep and Benchmarking

Model ReleasesDGX agent

arXiv:2602.03912v4 Announce Type: replace Abstract: This paper investigates the performance of Echo State Networks (ESNs) for univariate forecasting of monthly and quarterly time series from the M4 Fo

eMoT: evolving Memory-of-Thought via Symbolic Anchoring and Memory Corrosion

ResearchDGX agent

arXiv:2606.02054v1 Announce Type: new Abstract: While Large Language Models (LLMs) achieve impressive performance on multi-step reasoning tasks, their reliability is persistently hindered by critical

Enhancing BiGRU with a KAN Block for Legal Document Classification and Summarization

ApplicationsDGX agent

arXiv:2606.00116v1 Announce Type: cross Abstract: This study introduces a novel architecture of KAN-based BiGRU model for the task of classification and summarization of legal documents in a low-resou

EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams

Model ReleasesDGX agent

arXiv:2603.27223v2 Announce Type: replace-cross Abstract: We present EuraGovExam, a multilingual and multimodal benchmark sourced from real-world civil service examinations across five representative

FLaG: Fine-Grained Latent Grouping for Hallucination Detection

ResearchDGX agent

arXiv:2606.00301v1 Announce Type: new Abstract: Hallucinations in large language models (LLMs) arise from heterogeneous failure mechanisms, making reliable detection difficult for any single global un

Flow Matching for Convective-Scale Precipitation Downscaling

Model ReleasesDGX agent

arXiv:2606.00281v1 Announce Type: cross Abstract: Generative machine learning is an increasingly important complement to dynamical downscaling for producing high-resolution precipitation projections,

Flux klein 9b Comic Character Lora?

Local AiDGX agent

This post discusses creating or using a LoRA (Low-Rank Adaptation) model compatible with Flux Klein 9b, an AI image generation model, specifically for generating comic book-style characters. The discu

From Evaluation to Design: Using Potential Energy Surface Smoothness Metrics to Guide Machine Learning Interatomic Potential Architectures

Model ReleasesDGX agent

arXiv:2602.04861v2 Announce Type: replace-cross Abstract: Machine Learning Interatomic Potentials (MLIPs) sometimes fail to reproduce the physical smoothness of the quantum potential energy surface (P

From Outliers to Errors: Auditing Pali-to-English LLM Translations with Multi-Reference Adjudication

Model ReleasesDGX agent

arXiv:2606.01136v1 Announce Type: new Abstract: Single-score translation metrics can conflate legitimate variation with error, a problem especially acute for classical languages where multiple defensi

GABI: Geometry-Aware Boundary Integration for Spacecraft Segmentation

Model ReleasesDGX agent

arXiv:2606.00886v1 Announce Type: new Abstract: Accurate segmentation is crucial for autonomous spacecraft, as it directly affects downstream tasks related to 3D situational awareness. The harsh illum

GateKD: Confidence-Gated Closed-Loop Distillation for Robust Reasoning

ResearchDGX agent

arXiv:2605.13136v2 Announce Type: replace Abstract: Distilling multi-step reasoning abilities from large language models (LLMs) into compact student models remains challenging due to noisy rationales,

Global PIQA: Evaluating Commonsense Reasoning Across 100+ Languages and Cultures

Model ReleasesDGX agent

arXiv:2510.24081v2 Announce Type: replace Abstract: To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and

← Previous
1…400401402403404…1053
Next →