AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

87,042Total entries
1Added by human
87,041Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,509 results
7 Jul 2026

Auto: The AGI Compiler

Model ReleasesDGX agent

arXiv:2607.04542v1 Announce Type: cross Abstract: Every LLM agent run re-derives its behavior token by token on a frontier model: brilliant, expensive, slow, and unbounded. We present Auto, a compiler

Back to Basics: Improving Molecular Understanding in LLMs via SMILES-Graph Translation

Model ReleasesDGX agent

arXiv:2607.03007v1 Announce Type: cross Abstract: Recent advances in molecular large language models have led to strong performance on molecular understanding and generation tasks, yet these gains oft

Beyond Modality Fusion: Deep Ensembles for Multimodal Classification

Model ReleasesDGX agent

arXiv:2607.05019v1 Announce Type: cross Abstract: In multimodal classification, late-fusion approaches classify concatenated modality-specific features extracted by unimodal neural networks. When moda

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Beyond the Need for Speed: Energy-Aware Code Generation via Simulation-Guided Reinforcement Learning

ApplicationsDGX agent

arXiv:2607.04577v1 Announce Type: new Abstract: Code models strictly prioritize functional correctness, leaving software energy efficiency as an unoptimized byproduct. Training models to generate ener

CausalGame: Benchmarking Causal Thinking of LLM Agents in Games

Model ReleasesDGX agent

arXiv:2607.04293v1 Announce Type: cross Abstract: Building AI Scientist agents with Large Language Models (LLMs) has recently attracted growing attention. Since scientific discovery fundamentally reli

CineMobile: On-Device Image-to-Video Diffusion for Cinematic Camera Motion Generation

Model ReleasesDGX agent

arXiv:2607.03803v1 Announce Type: cross Abstract: The growing demand for image-to-video creation on mobile devices has increasingly focused on cinematic motion effects like bullet time, dolly zoom, sl

Code Benchmarks Should Prioritize Rigor, Reliability, and Reproducibility

Model ReleasesDGX agent

arXiv:2501.10711v5 Announce Type: replace-cross Abstract: Code-related benchmarks play a critical role in evaluating large language models (LLMs), yet their quality fundamentally shapes how the commun

Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

SafetyDGX agent

arXiv:2508.10031v2 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and

CritiqueDriveVLM: From Verifier-Guided Reinforcement Learning to Latent Thought Distillation for Autonomous Driving

Model ReleasesDGX agent

arXiv:2607.04179v1 Announce Type: cross Abstract: End-to-end Vision-Language Models (VLMs) show immense potential in autonomous driving. However, standard Supervised Fine-Tuning (SFT) often suffers fr

Data Driven Optimization of GPU efficiency for Distributed LLM-Adapter Serving

HardwareDGX agent

arXiv:2602.24044v2 Announce Type: replace-cross Abstract: Large Language Model (LLM) adapters enable low-cost model specialization, but introduce complex caching and scheduling challenges in distribut

Discrete distributions are learnable from metastable samples

Model ReleasesDGX agent

arXiv:2410.13800v4 Announce Type: replace-cross Abstract: Physically motivated stochastic dynamics are widely used to sample from high-dimensional distributions. However, such samplers often get trapp

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation

Model ReleasesDGX agent

arXiv:2607.05147v1 Announce Type: new Abstract: Speculative decoding accelerates Large Language Model (LLM) inference by decoupling draft generation from target verification. While recent parallel dra

Dual-Adaptive SAM3: Hierarchical Routing over Low-Rank Expert Layers for Parameter-Efficient Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.02571v1 Announce Type: new Abstract: The Segment Anything Model with Concepts (SAM3) heralds a new paradigm for open-vocabulary segmentation through natural language interaction, offering s

FUSE: FK-Steered Multi-Modal Flow Matching for Efficient Simulation-Based Posterior Estimation

Model ReleasesDGX agent

arXiv:2607.05252v1 Announce Type: new Abstract: Simulation-Based Inference (SBI) is critical for scientific discovery, with generative models offering a promising path toward efficient inference. Howe

FuseMamba-VD: Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection

Model ReleasesDGX agent

arXiv:2506.03162v3 Announce Type: replace-cross Abstract: The rapid proliferation of surveillance cameras has increased the demand for automated violence detection. While CNNs and Transformers have sh

Governed MCP: Kernel-Level Tool Governance for AI Agents via Logit-Based Safety Primitives

Model ReleasesDGX agent

arXiv:2604.16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP). These tool calls are the age

Grokking Is Conditional and Fragile: A Fully-Tractable, Multi-Seed Study at 12K Parameters

Model ReleasesDGX agent

arXiv:2607.05104v1 Announce Type: cross Abstract: Grokking -- the delayed onset of generalization long after a network has fit its training set - -is usually studied in models too large to read comple

GuideMe: Multi-Domain Task Guidance and Intervention in Streaming Video

Model ReleasesDGX agent

arXiv:2607.02991v1 Announce Type: new Abstract: While multimodal Large Language Models (MLLMs) excel at offline video understanding, an interesting question of how far they are from serving as a real-

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation

Model ReleasesDGX agent

arXiv:2607.04344v1 Announce Type: cross Abstract: While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like ocular sur

K9-Bench: Evaluating Multimodal LLMs on Canine-Centric Videos

Model ReleasesDGX agent

arXiv:2607.02680v1 Announce Type: cross Abstract: MLLMs have shown strong zero-shot capabilities across diverse inputs such as across images, video, audio, and text. A crucial, yet underexplored, appl

Learning When to Attend: Conditional Memory Access for Long-Context LLMs

Model ReleasesDGX agent

arXiv:2603.17484v2 Announce Type: replace Abstract: Language models struggle to generalize beyond pretraining context lengths, limiting long-horizon reasoning and retrieval. Continued pretraining on l

Legible-by-Construction: Attention and End-to-End Transformers

Model ReleasesDGX agent

arXiv:2607.04319v1 Announce Type: new Abstract: A companion paper showed that a transformer's feed-forward layer can be rebuilt from explicit fuzzy set operations - intersection, set-difference, and a

Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption

HardwareDGX agent

arXiv:2607.04553v1 Announce Type: cross Abstract: We present a bidirectional framework for estimating the energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models from architect

LLM-Based Test Oracles: Source-of-Authority Taxonomy -- A Systematic Literature Review

Model ReleasesDGX agent

arXiv:2607.05031v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to produce test oracles, the part of a test that decides whether observed behavior is correct. Yet

Localized LoRA-MoE: Block-wise Low-Rank Experts With Adaptive Routing

Model ReleasesDGX agent

arXiv:2607.05114v1 Announce Type: cross Abstract: Large Language Models (LLMs) and high-dimensional perception networks increasingly rely on parameter-efficient fine-tuning (PEFT) to adapt to diverse

Measuring the Robustness of Audio Deepfake Detection under Real-World Corruption

ApplicationsDGX agent

arXiv:2503.17577v2 Announce Type: replace-cross Abstract: Deepfakes have emerged as a widespread and rapidly escalating concern in generative AI, spanning images, audio, and videos. Among these, audio

MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources

ApplicationsDGX agent

arXiv:2601.22054v2 Announce Type: replace-cross Abstract: Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging du

MV-Forcing: Long Multi-View Video Generation via 4D-Grounded Spatio-Temporal Self-Forcing

SafetyDGX agent

arXiv:2607.05376v1 Announce Type: new Abstract: Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis

Natural Language Camera Movement Understanding

Model ReleasesDGX agent

arXiv:2607.03043v1 Announce Type: new Abstract: Understanding camera movement in natural language is critical for training and evaluating video generation models, among other applications. However, we

Not All Refusals Are Equal: How Safety Alignment Fails Cybersecurity at Scale

Model ReleasesDGX agent

arXiv:2607.02714v1 Announce Type: cross Abstract: There is no doubt that safety alignment is an essential step in LLM training. However, conceptually it does not distinguish between various domains an

Polarity Detection of Sustainable Development Goals in News Text

Model ReleasesDGX agent

arXiv:2509.19833v4 Announce Type: replace-cross Abstract: The United Nations' Sustainable Development Goals (SDGs) provide a globally recognised framework for addressing major societal, environmental,

Probe, Don't Prompt: A Hidden-State Probe for Metadata Filtering in Multi-Meta-RAG

Model ReleasesDGX agent

arXiv:2607.03929v1 Announce Type: cross Abstract: Multi-Meta-RAG improves retrieval for multi-hop question answering by filtering a vector store on metadata (the news source) that it extracts from eac

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs

Model ReleasesDGX agent

arXiv:2602.20629v3 Announce Type: replace Abstract: As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated ev

Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents

Model ReleasesDGX agent

arXiv:2607.03968v1 Announce Type: cross Abstract: Large language models are increasingly deployed as IDE-integrated coding agents that decompose tasks, generate and edit files, run code, and refine ou

Restricted Bernoulli Matrix Factorization: Balancing the trade-off between prediction accuracy and coverage in classification based collaborative filtering

ResearchDGX agent

arXiv:2210.10619v3 Announce Type: replace-cross Abstract: Reliability measures associated with the prediction of the machine learning models are critical to strengthening user confidence in artificial

SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding

Model ReleasesDGX agent

arXiv:2607.04017v1 Announce Type: new Abstract: Human object interaction (HOI), gaze pattern, and their anticipation are intricately linked, providing valuable insights into cognitive processes, inten

Seduced by the Narrative: Assessing Rule Adherence in Semi-Open Textual Sandboxes

Model ReleasesDGX agent

arXiv:2607.02802v1 Announce Type: cross Abstract: As LLMs are increasingly deployed as autonomous adjudicators in semi-open textual game environments, robust rule adherence becomes critical when user

SNR-Adaptive Unified Diffusion for Multi-Task Medical Image Segmentation

Model ReleasesDGX agent

arXiv:2607.03103v1 Announce Type: cross Abstract: Clinical cardiac imaging pipelines currently deploy separate models for each dataset and modality, incurring redundant training costs and precluding k

Solve the Missing First Step: Can VLMs Standardize Raw Heterogeneous Medical Data?

Model ReleasesDGX agent

arXiv:2607.04694v1 Announce Type: new Abstract: As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over giv

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

Model ReleasesDGX agent

arXiv:2511.07403v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in vision-language tasks, but continue to struggle with spatial rea

Structured Prompting and Automated Evaluation in Fixed Synthetic Japanese-Language Counseling Dialogues

Model ReleasesDGX agent

arXiv:2507.02950v3 Announce Type: replace-cross Abstract: Large language models (LLMs) may support counseling training, yet evidence from Japanese-language interactions and automated quality ratings r

StructuredEdit: Constraint-Aware Graphic Design Editing via Differentiable Parameter Propagation

Model ReleasesDGX agent

arXiv:2607.04612v1 Announce Type: cross Abstract: Graphic design editing requires precise manipulation of typography, layout, and visual hierarchy under strict design constraints. Following the introd

The 'I Don't Know' Filter: Enhancing Agentic Reliability in Function Calling

AgentsDGX agent

arXiv:2607.04034v1 Announce Type: cross Abstract: The language models that underpin agents have seen a rapid rise in performance on function calling benchmarks. However, the metrics used in the traini

The Method of Gaps: Exact Expressions for the Generalization Error of Supervised Learning Algorithms

Model ReleasesDGX agent

arXiv:2411.12030v3 Announce Type: replace Abstract: In this paper, the method of gaps, a technique for deriving closed-form expressions in terms of information measures for the generalization error of

Tile-Level Activation Overlap for Efficient LLM Inference

Model ReleasesDGX agent

arXiv:2607.02521v1 Announce Type: cross Abstract: SwiGLU is the dominant MLP activation in modern large language models, yet its intermediate tensor materialization costs 9-37% of MLP execution time.

TimeThink: Reasoning with Time for Video LLMs

ResearchDGX agent

arXiv:2607.05089v1 Announce Type: new Abstract: Video reasoning requires models to identify and verify temporally localized evidence within long video sequences. Recent Video Large Language Models (Vi

Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation

ResearchDGX agent

arXiv:2607.02593v1 Announce Type: cross Abstract: While knowledge distillation (KD) is widely adopted for training lightweight models by leveraging supervision from larger teacher models, relying sole

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker

Model ReleasesDGX agent

arXiv:2605.25706v2 Announce Type: replace Abstract: Referring expression comprehension (REC) aims to localize a target object within an image based on a given expression. Although recent advances in v

Transformers with Physics-Informed Encodings and Simulation-Based Inference for Robust Detection of Eccentric Binary Black Holes in Pulsar Timing Array Data

Model ReleasesDGX agent

arXiv:2607.03904v1 Announce Type: new Abstract: Pulsar timing arrays (PTAs) provide a unique window into nanohertz gravitational waves (GWs), but extracting astrophysical parameters from noisy, long-b

Transplanting, inverting, and preventing a misalignment persona: method-conditional emergent misalignment in Qwen2.5

ApplicationsDGX agent

arXiv:2607.04510v1 Announce Type: cross Abstract: Emergent misalignment (EM) -- the broad misbehaviour a language model acquires after fine-tuning on narrow harmful data -- is mediated in Qwen2.5 mode

6 Jul 2026

By watching the J-space, we can see Claude silently perform reasoning steps in its head—noticing bugs in code, identifying images, and more.

Model ReleasesDGX agent

This post discusses how observing Claude's internal activation patterns (J-space) reveals the model's hidden reasoning processes, such as detecting code bugs and analyzing images, even when these step

Reporting benchmark results as a scalar number, e.g. '75% on XYZ' is completely meaningless at this point. You should always report efficien…

Model ReleasesDGX agent

Reporting benchmark results as a single scalar percentage is insufficient for meaningful evaluation of model performance. Comprehensive benchmark reporting should include efficiency metrics alongside

Shift into high gear with agents: Securing the software-defined vehicle

Model ReleasesDGX agent

The automotive industry is at a pivotal crossroads as it hits the gas on adopting new technology. The era of the traditional connected vehicle has shifted into the age of the software-defined vehicle

3 Jul 2026

A global predicted-fMRI drive signal from TRIBE does not predict YouTube replay heatmaps

Model ReleasesDGX agent

arXiv:2607.01400v1 Announce Type: cross Abstract: Deep multimodal brain-encoding models now predict fMRI responses to naturalistic video with high accuracy. Whether their predicted neural signals also

AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG

Model ReleasesDGX agent

arXiv:2602.19127v2 Announce Type: replace Abstract: With the rapid advancement of agent-based methods in recent years, Agentic RAG has undoubtedly become an important research direction. Multi-hop rea

AIriskEval-edu: New Dataset for Risk Assessment in AI-mediated K-12 Educational Explanations

Model ReleasesDGX agent

arXiv:2607.01934v1 Announce Type: cross Abstract: This work introduces AIriskEval-edu-db2, a new dataset designed to train and evaluate auditors based on LLMs for an explainable pedagogical risk asses

Assessing VLM Reliability for Medical Image Quality Evaluation Under Corruption and Bias

Model ReleasesDGX agent

arXiv:2607.01973v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly applied in medical tasks such as pathology description, report generation, and visual question answerin

Autorelevance function and other feature relevance measures for univariate time series

ResearchDGX agent

arXiv:2607.01959v1 Announce Type: cross Abstract: We propose a model agnostic methodology to measure lag relevance in machine learning forecasting models applied to univariate time series. Particularl

Benchmarking Federated Learning and Knowledge Distillation for Point Cloud Classification

Model ReleasesDGX agent

arXiv:2607.01272v1 Announce Type: cross Abstract: Deploying 3D point cloud analysis in privacy-sensitive, resource-constrained settings faces two barriers: data cannot be centralized, and models must

BOUNDARY_SYNC: Measuring Communication-Induced Representational Coupling in Multi-Agent LLM Systems

Model ReleasesDGX agent

arXiv:2607.01600v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed as communicating agents, does inter-agent communication cause outputs to converge? We introduce BOUNDARY_

← Previous
1…336337338339340…1042
Next →