AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

87,042Total entries
1Added by human
87,041Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,509 results
23 Jun 2026

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

TutorialsDGX agent

arXiv:2606.22565v1 Announce Type: cross Abstract: Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LLMs) by eliciting step-by-step thi

Mimic Human Cognition, Master Multi-Image Reasoning: A Meta-Action Framework for Enhanced Visual Understanding

ResearchDGX agent

arXiv:2601.07298v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) excel at single-image understanding, they exhibit significantly degraded performance in multi-image r

Mirage: a Clean-Label Backdoor against LiDAR 3D Object Detection

Model ReleasesDGX agent

arXiv:2606.20752v1 Announce Type: new Abstract: Deep neural network-based LiDAR 3D object detection serves as a critical perception component in safety-critical autonomous systems. However, recent stu

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Multigrid Training for Molecular Generation using Graph Neural Networks

Model ReleasesDGX agent

arXiv:2606.22377v1 Announce Type: new Abstract: Deep learning has demonstrated significant success for modeling biochemical molecular systems, where inputs are commonly represented as graphs or 3D gri

OmniV2X: A Generative Foundation Planner for Efficient End-to-End Cooperative Driving

AgentsDGX agent

arXiv:2606.21165v1 Announce Type: new Abstract: We present OmniV2X, a generative foundation model for vehicle-to-everything (V2X) cooperative driving. The model directly interprets independent context

ORBIT: Training-Free Multi-Attribute Behavioral Steering via Orthogonal Subspace Rotation

Model ReleasesDGX agent

arXiv:2606.22357v1 Announce Type: cross Abstract: Language models are widely used in assistant settings, where controlling behavioral attributes is often essential. Activation steering modifies hidden

PROTON: Prototype-Based Test-Time Online OOD Detection for Medical VLMs

Model ReleasesDGX agent

arXiv:2606.20913v1 Announce Type: new Abstract: Medical vision-language models (VLMs) enable zero-shot clinical image classification, yet reliably detecting out-of-distribution (OOD) inputs at deploym

Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild

Model ReleasesDGX agent

arXiv:2603.04205v2 Announce Type: replace Abstract: While Vision-Language Models (VLMs) achieve near-perfect scores on digital document benchmarks like OmniDocBench, their performance in the unpredict

Revisiting the Neural Tangent Kernel: the role of large width and depth

Model ReleasesDGX agent

arXiv:2511.07272v2 Announce Type: replace Abstract: Overparameterized fully-connected neural networks have been shown to behave like kernel models when trained with gradient descent, assuming standard

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

Model ReleasesDGX agent

arXiv:2606.23221v1 Announce Type: new Abstract: Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. Howev

Sakana Fugu Technical Report

AgentsDGX agent

arXiv:2606.21228v1 Announce Type: new Abstract: The capabilities of frontier Large Language Models (LLMs) continue to advance, with different providers increasingly specializing in distinct domains. T

SATURN: Symbolic Spatial Reasoning for Multi-Perspective Grounding

Model ReleasesDGX agent

arXiv:2606.22694v1 Announce Type: new Abstract: Vision-Language Models (VLMs) remain unreliable when spatial reasoning requires composing relations whose meanings depend on frames of reference. Existi

Scaling Linear Mode Connectivity and Merging to Billion Parameter Pretrained Transformers

Model ReleasesDGX agent

arXiv:2606.23607v1 Announce Type: new Abstract: Linear mode connectivity (LMC) provides a promising foundation for understanding and merging independently trained neural networks, but existing methods

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning

Model ReleasesDGX agent

arXiv:2606.22873v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed in consumer, medical, financial, and enterprise applications. This broad deployment expands the

Specialize Roles, Mix Deployments: Pushing the Cost-Accuracy Frontier of LLM Agent Teams

Model ReleasesDGX agent

arXiv:2606.20629v1 Announce Type: cross Abstract: LLM agents are increasingly deployed as multi-role teams, where tasks are divided across specialized roles such as planner, executor, and verifier. In

T-IMPACT: A Severity-Aware Benchmark for Contextual Image-Text Manipulation

Model ReleasesDGX agent

arXiv:2606.22339v1 Announce Type: new Abstract: Recent advances in vision-language models and generative editing systems have made it increasingly easy to produce persuasive multimodal misinformation

Temporal-Spectral Alignment with Frequency Adaptation for Source-Free Time-Series Adaptation

Model ReleasesDGX agent

arXiv:2606.23120v1 Announce Type: new Abstract: The goal of source-free domain adaptation (SFDA) for time-series data is to transfer knowledge from a pre-trained source model to an unlabeled target do

The Alignment Problem in Constrained Code Generation

SafetyDGX agent

arXiv:2606.21619v1 Announce Type: cross Abstract: Large Language Models (LLMs) have demonstrated strong capabilities in code generation, but their outputs frequently contain syntax or type errors that

Topological Out-of-Domain Generalization in Dynamical Systems Reconstruction

Model ReleasesDGX agent

arXiv:2606.22969v1 Announce Type: new Abstract: Predicting the behavior of dynamical systems (DS) beyond the dynamical and parameter regimes observed in training is a pivotal and essentially unresolve

Towards Robust Personalized Federated Learning: Vulnerability Assessment and Defense Co-Design

Model ReleasesDGX agent

arXiv:2606.22782v1 Announce Type: new Abstract: The proliferation of IoT devices has fueled distributed edge systems to collect vast amounts of sensitive data, creating fertile ground for on-device ma

TROPT: An Open Framework for Unifying and Advancing Discrete Text Optimization

SafetyDGX agent

arXiv:2606.23496v1 Announce Type: new Abstract: Discrete text-trigger optimization -- searching for text sequences that, when ingested by a model, steer it toward a specified objective -- underpins mo

Using predictive multiplicity to measure individual performance within the AI Act

SafetyDGX agent

arXiv:2602.11944v2 Announce Type: replace Abstract: When building AI systems for decision support, one often encounters the phenomenon of predictive multiplicity: a single best model does not exist; i

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

Model ReleasesDGX agent

arXiv:2606.22079v1 Announce Type: cross Abstract: Web data curation has been widely studied for decoder Large Language Model (LLM) pretraining. Encoders for dense-terminology domains such as medicine,

22 Jun 2026

How does it work? Sakana Fugu is itself an LLM, trained to call various LLMs in an agent pool, including instances of itself recursively. Fu…

AgentsDGX agent

How does it work? Sakana Fugu is itself an LLM, trained to call various LLMs in an agent pool, including instances of itself recursively. Fugu dynamically orchestrates the world's best models to tackl

Sakana Fugu Ultra now available on AI Gateway

ToolsDGX agent

Sakana Fugu Ultra, a new AI model, is now available through Vercel's AI Gateway, expanding the selection of models developers can access via the platform. This addition allows users to integrate Sakan

21 Jun 2026

An hour in and first impression is definitely that GLM is really solid (very easy to set up on @FireworksAI_HQ, props to them for that, took…

Model ReleasesDGX agent

A user shares positive early impressions of GLM (likely a language model), praising its solid performance and ease of setup on Fireworks AI's platform. The post highlights Fireworks AI's developer exp

20 Jun 2026

I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on RTX PRO 6000, H100, B20…

Model ReleasesDGX agent

I have some very big news... KernelBench-Hard with H100 and B200 (single gpu results) AND KernelBench-Mega tested on RTX PRO 6000, H100, B200 is finally out! Starting with Mega, each of models wrote a

11 Jun 2026

ALIGNBEAM : Inference-Time Alignment Transfer via Cross-Vocabulary Logit Mixing

SafetyDGX agent

arXiv:2606.12342v1 Announce Type: cross Abstract: Domain fine-tuning degrades the safety of large language models: fine-tuned specialists readily comply with harmful prompts framed in domain language.

Architecture-Aware Reinforcement Learning Makes Sliding-Window Attention Competitive in Math Reasoning

SafetyDGX agent

arXiv:2606.11634v1 Announce Type: new Abstract: The rapid progress of reasoning and agentic large language models (LLMs) has increased the demand for long-context inference, but self-attention (SA) sc

Beyond Third-Person Audits: Situated Interaction Auditing for User-Centered LLM Bias Research

SafetyDGX agent

arXiv:2606.12247v1 Announce Type: cross Abstract: Research on bias in large language models (LLMs) has predominantly focused on third-person audits, which study how models represent or evaluate demogr

Bootstrapped Monitoring: Leveraging Transparent Reasoning to Oversee Stronger AI Agents

AgentsDGX agent

arXiv:2606.11998v1 Announce Type: new Abstract: Trusted monitoring is a cornerstone of AI control. However, as frontier models grow more capable, the increasing capabilities gap between trusted and un

Claw-SWE-Bench: A Benchmark for Evaluating OpenClaw-style Agent Harnesses on Coding Tasks

Model ReleasesDGX agent

arXiv:2606.12344v1 Announce Type: cross Abstract: General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-ben

Dummy Backdoor as a Defense: Removing Unknown Backdoors via Shared Internal Mechanisms for Generative LLMs

SafetyDGX agent

arXiv:2606.11648v1 Announce Type: cross Abstract: Backdoor attacks pose a serious threat to the safety and reliability of Large Language Models (LLMs), as they cause models to behave normally on clean

Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training

Model ReleasesDGX agent

arXiv:2606.11854v1 Announce Type: cross Abstract: There are two main Parameter-Efficient Fine-Tuning (PEFT) techniques for Large Language Models (LLMs). While Low-Rank Adaptation (LoRA) introduces add

FronTalk: Benchmarking Front-End Development as Conversational Code Generation with Multi-Modal Feedback

Model ReleasesDGX agent

arXiv:2601.04203v2 Announce Type: replace Abstract: We present FronTalk, a benchmark for front-end code generation that pioneers the study of a unique interaction dynamic: conversational code generati

How an astrophysicist uses Codex to help simulate black holes

Model ReleasesDGX agent

An astrophysicist leverages OpenAI's Codex AI model to accelerate the development of code for simulating black hole physics and behavior. Codex assists in generating complex scientific code more effic

Improving Detection of Rare Nodes in Hierarchical Multi-Label Learning

Model ReleasesDGX agent

arXiv:2602.08986v2 Announce Type: replace-cross Abstract: In hierarchical multi-label classification, a persistent challenge is enabling model predictions to reach deeper levels of the hierarchy for m

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends

Model ReleasesDGX agent

arXiv:2606.12207v1 Announce Type: cross Abstract: Embodied intelligence now spans navigation, household assistance, manipulation, autonomous driving, aerial agents, and multimodal large-model control.

Lung-SRAD: Spectral-Aware Regularized Audio DASS with Dual-Axis Patch-Mix Contrastive Learning for Respiratory Sound Classification

Model ReleasesDGX agent

arXiv:2606.11922v1 Announce Type: cross Abstract: Recent respiratory sound classification (RSC) studies largely rely on CLS-token driven self-attention architectures such as the Audio Spectrogram Tran

MARIC: Multi-Agent Reasoning for Image Classification

Model ReleasesDGX agent

arXiv:2509.14860v2 Announce Type: replace-cross Abstract: Image classification has traditionally relied on parameter-intensive model training, requiring large-scale annotated datasets and extensive fi

MobilityBench: A Benchmark for Evaluating Route-Planning Agents in Real-World Mobility Scenarios

Model ReleasesDGX agent

arXiv:2602.22638v2 Announce Type: replace Abstract: Route-planning agents powered by large language models (LLMs) have emerged as a promising paradigm for supporting everyday human mobility through na

Noise-Aware Framework for Correcting Corrupted Labels

ApplicationsDGX agent

arXiv:2606.11695v1 Announce Type: cross Abstract: High-quality labeled data is essential for training reliable ML/DL models. However, real-world datasets often contain a considerable proportion of cor

ProGRank: Probe-Gradient Reranking to Defend Dense-Retriever RAG from Corpus Poisoning

Model ReleasesDGX agent

arXiv:2603.22934v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) improves large language model applications by grounding generation in retrieved evidence, but also introduces c

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding

Model ReleasesDGX agent

arXiv:2606.12125v1 Announce Type: new Abstract: Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames

Reassessing High-Performing LLMs on Polish Medical Exams: True Competence or Bias-Driven Performance?

Model ReleasesDGX agent

arXiv:2606.12250v1 Announce Type: new Abstract: Large language models (LLMs) in medicine are mainly evaluated using multiple-choice question answering (MCQA), which can overestimate real clinical abil

RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

SafetyDGX agent

arXiv:2606.11709v1 Announce Type: cross Abstract: On-policy self-distillation (OPSD) provides dense, token-level supervision for reasoning models by aligning a model's own distribution with the distri

SDQM: Synthetic Data Quality Metric for Object Detection Dataset Evaluation

ResearchDGX agent

arXiv:2510.06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data. The scarcity of large-scale, well-annotated datasets poses signif

Sparsified Kolmogorov-Arnold Networks for Interpretable Quantum State Tomography

Model ReleasesDGX agent

arXiv:2606.11814v1 Announce Type: cross Abstract: Machine-learning approaches to quantum state tomography can achieve high reconstruction fidelity, but the physical structure used by the trained model

System Report for CCL25-Eval Task 5: New Dataset and LoRA-Fine-Tuned Qwen2.5

Model ReleasesDGX agent

arXiv:2606.12392v1 Announce Type: cross Abstract: Recently, large language models (LLMs) have achieved promising progress in the fields of classical Chinese translation and the generation of classical

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video

Model ReleasesDGX agent

arXiv:2512.11393v2 Announce Type: replace Abstract: Humans can intuitively parallelise complex activities, but can a model predict this from observing a single person? Given one egocentric video, we i

Toward Trustworthy AI: Multi-Target Adversarial Attacks and Robust Defenses for Continuous Data Summarization

Model ReleasesDGX agent

arXiv:2606.11804v1 Announce Type: new Abstract: Trustworthy AI requires reliable data-processing pipelines, not only robust downstream predictive models. As an upstream component, data summarization d

10 Jun 2026

AdaGC: Enhancing LLM Pretraining Stability via Adaptive Gradient Clipping

Model ReleasesDGX agent

arXiv:2502.11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining. While previous research has attempted to identify the root cause

AgenticNav: Zero-Shot Vision-and-Language Navigation as a Tool-Calling Harness

Model ReleasesDGX agent

arXiv:2606.10577v1 Announce Type: new Abstract: Zero-shot vision-and-language navigation in continuous environments (VLN-CE) has recently become feasible with large vision-language models (VLMs). Howe

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

ResearchDGX agent

arXiv:2606.10657v1 Announce Type: new Abstract: Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes t

Benchmarking Knowledge Editing using Logical Rules

Model ReleasesDGX agent

arXiv:2606.10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge. However, retraining LLM

ChartLens: A Dual-Branch Framework for Chart Data Correction and Factual Summary Refinement

Model ReleasesDGX agent

arXiv:2606.10640v1 Announce Type: new Abstract: In this report, we present our champion solution for the DataMFM Challenge Track 2: Chart Understanding. This track requires models to recover structure

CodeAlchemy: Synthetic Code Rewriting at Scale

Model ReleasesDGX agent

arXiv:2606.10087v1 Announce Type: new Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative f

DiffusionGemma: 4x faster text generation

Model ReleasesDGX agent

DiffusionGemma is an experimental open model from Google DeepMind that uses text diffusion for exceptionally fast generation, moving beyond sequential token-by-token processing to generate entire bloc

Evaluating Research-Level Math Proofs via Strict Step-Level Verification

Model ReleasesDGX agent

arXiv:2606.10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs. Standard global evaluation approaches suffer from 'context poiso

If Claude Fable stops helping you, you'll never know

Model ReleasesDGX agent

If Claude Fable stops helping you, you'll never know Jonathon Ready highlights one of the more eyebrow-raising details from the 319 page system card for Fable 5 and Mythos 5. Here's a longer excerpt,

← Previous
1…341342343344345…1042
Next →