AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,103 results
2 Jul 2026

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

Model ReleasesDGX agent

arXiv:2607.01002v1 Announce Type: cross Abstract: In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pastin

Low Perplexity is Repetition: A One-Dimensional Self-Conditioning Attractor in Continuous Diffusion LMs

ResearchDGX agent

arXiv:2607.00588v1 Announce Type: new Abstract: Continuous diffusion language models such as ELF report record-low generative perplexity (Gen-PPL). We find a catch: these models repeat far more than h

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

Model ReleasesDGX agent

arXiv:2510.24434v3 Announce Type: replace Abstract: The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resource linguistic settings due to a lack of high-quali

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Quantum vs. Classical Machine Learning: A Unified Empirical Comparison

Model ReleasesDGX agent

arXiv:2607.01197v1 Announce Type: new Abstract: Quantum computing has emerged as a promising computational paradigm for machine learning (ML), with the potential to offer computational advantages over

Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval

SafetyDGX agent

arXiv:2607.00448v1 Announce Type: cross Abstract: The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards for training

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

Model ReleasesDGX agent

arXiv:2607.00816v1 Announce Type: new Abstract: High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when t

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

Model ReleasesDGX agent

arXiv:2511.17731v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential

WorkBench Revisited: Workplace Agents Two Years On

Model ReleasesDGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

Model ReleasesDGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

1 Jul 2026

Addressing Over-Refusal in LLMs with Competing Rewards

SafetyDGX agent

arXiv:2606.31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Tho

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

Model ReleasesDGX agent

arXiv:2606.31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) model

Efficient Public Verification of Private ML via Regularization

Model ReleasesDGX agent

arXiv:2512.04008v2 Announce Type: replace Abstract: Training with differential privacy (DP) guarantees dataset members that they cannot be identified by users of the released model. However, those dat

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

Model ReleasesDGX agent

Hugging Face and Cerebras have collaborated to integrate Google's Gemma 4 model with real-time voice AI capabilities, enabling faster speech processing and voice interactions. This integration likely

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

Model ReleasesDGX agent

arXiv:2606.31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant atten

Loved the chat between @trq212 @_catwu @simonw at AI Eng summit. My top 13 takeaways from their session -> 1. Engineers should become better…

Model ReleasesDGX agent

Loved the chat between @trq212 @_catwu @simonw at AI Eng summit. My top 13 takeaways from their session -> 1. Engineers should become better at product/business sense. 2. Don't worry about major rewri

MAPE: Defending Against Transferable Adversarial Attacks Using Multi-Source Adversarial Perturbations Elimination

ResearchDGX agent

arXiv:2606.31378v1 Announce Type: new Abstract: Neural networks are vulnerable to meticulously crafted adversarial examples, leading to high-confidence misclassifications in image classification tasks

Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

Model ReleasesDGX agent

arXiv:2606.31543v1 Announce Type: new Abstract: Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

Model ReleasesDGX agent

arXiv:2606.23672v2 Announce Type: replace Abstract: This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In this tas

Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer

Model ReleasesDGX agent

arXiv:2606.31574v1 Announce Type: cross Abstract: Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service life of fusi

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

Model ReleasesDGX agent

arXiv:2606.31704v1 Announce Type: new Abstract: The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance dispari

30 Jun 2026

A Machine-Verified Proof of a Quantum-Optimization Conjecture

Model ReleasesDGX agent

arXiv:2606.29687v1 Announce Type: cross Abstract: We report a machine-verified resolution of a problem open for over a decade in quantum optimization: the Farhi, Goldstone and Gutmann (FGG) conjecture

A Trainable-by-Parts Operator Learning Framework: Bridging DeepONet and Karhunen-Loeve Expansions for Large-Scale Applications

HardwareDGX agent

arXiv:2606.28519v1 Announce Type: new Abstract: Training operator-learning models for large-scale problems governed by partial differential equations (PDEs) is challenging due to the curse of dimensio

Adaptive Financial Transformer with Regime-Gated Attention for Stock Return Prediction

Model ReleasesDGX agent

arXiv:2606.29347v1 Announce Type: cross Abstract: Adaptive Financial Transformer (AFT) is proposed for stock return prediction under non-stationary financial markets. The model incorporates a Market R

Agent Safety Is Action Alignment

SafetyDGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching

ResearchDGX agent

arXiv:2509.15942v3 Announce Type: replace-cross Abstract: Internal variability is a dominant contributor to the uncertainty of predictions at the interannual to decadal timescale. A typical approach t

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

Model ReleasesDGX agent

arXiv:2606.30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

Model ReleasesDGX agent

arXiv:2606.29445v1 Announce Type: cross Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarka

Counterfactual Residual Data Augmentation for Regression

Model ReleasesDGX agent

arXiv:2606.28460v1 Announce Type: cross Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations. Inspir

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

Model ReleasesDGX agent

arXiv:2606.30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while th

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

HardwareDGX agent

arXiv:2606.29082v1 Announce Type: new Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrate

Factorizable Normalizing Flows for parameter-dependent density morphing

Model ReleasesDGX agent

arXiv:2606.30489v1 Announce Type: cross Abstract: Normalizing Flows excel at modeling a single fixed density, yet many problems across the sciences, such as high energy physics, instead require modeli

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring

Model ReleasesDGX agent

arXiv:2606.30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated. We ask wh

Longitudinal Lesion Inpainting in Brain MRI via 3D Region Aware Diffusion

Model ReleasesDGX agent

arXiv:2603.05693v2 Announce Type: replace-cross Abstract: Accurate longitudinal analysis of brain MRI is often hindered by evolving lesions, which bias automated neuroimaging pipelines. While deep gen

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

Local AiDGX agent

arXiv:2606.16364v2 Announce Type: replace Abstract: LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a

Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain

ResearchDGX agent

arXiv:2606.28772v1 Announce Type: cross Abstract: Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this aggregation is

Memory-Managed Long-Context Attention: A Preliminary Study of Editable Request-Local Memory

Model ReleasesDGX agent

arXiv:2606.28876v1 Announce Type: new Abstract: Long-context language models often conflate two different goals: compressing history into an efficient state, and maintaining reliable long-term memory.

MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

Model ReleasesDGX agent

arXiv:2602.21608v2 Announce Type: replace Abstract: Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain s

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

Model ReleasesDGX agent

arXiv:2606.29717v1 Announce Type: cross Abstract: Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of work has pr

Probabilistic Approach to Black-Box Binary Optimization with Budget Constraints: Application to Sensor Placement

Model ReleasesDGX agent

arXiv:2406.05830v2 Announce Type: replace-cross Abstract: This paper presents a fully probabilistic approach for solving optimal experimental design problems under budget constraints. The experimental

Reliability-Prioritized Fine-Grained Generation in Multimodal Large

Model ReleasesDGX agent

arXiv:2606.29573v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theo

RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots

Model ReleasesDGX agent

arXiv:2606.28649v1 Announce Type: cross Abstract: We present RIPA, the first systematic multi-channel empirical study of prompt injection attacks delivered through the sensory pipeline of a ROS 2-base

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.30201v1 Announce Type: cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical

Singular Learning and Occam's Razor in Deep Monomial Networks

Model ReleasesDGX agent

arXiv:2606.28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical poi

SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis

Local AiDGX agent

arXiv:2606.29586v1 Announce Type: cross Abstract: Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recent

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

Model ReleasesDGX agent

arXiv:2606.29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks

StrucTab: A Structured Optimization Framework for Table Parsing

Model ReleasesDGX agent

arXiv:2606.29905v1 Announce Type: new Abstract: Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial l

Test-Time Detoxification without Training or Learning Anything

SafetyDGX agent

arXiv:2602.02498v2 Announce Type: replace-cross Abstract: Large language models can produce toxic or inappropriate text even for benign inputs, creating risks when deployed at scale. Detoxification is

Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding

Model ReleasesDGX agent

arXiv:2601.04693v2 Announce Type: replace Abstract: Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scar

Towards Engineering Scaling Laws with Pretraining Data Composition

ResearchDGX agent

arXiv:2606.19781v2 Announce Type: replace-cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established fo

Translating Natural Language to Strategic Temporal Specifications via LLMs

Model ReleasesDGX agent

arXiv:2606.30441v1 Announce Type: cross Abstract: A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS). However, writing corr

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

Model ReleasesDGX agent

arXiv:2606.28438v1 Announce Type: cross Abstract: Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality control. We st

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

Model ReleasesDGX agent

arXiv:2606.28661v1 Announce Type: cross Abstract: People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard question by sam

Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression

Model ReleasesDGX agent

arXiv:2606.29712v1 Announce Type: new Abstract: Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and

29 Jun 2026

CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching

Model ReleasesDGX agent

arXiv:2602.20094v2 Announce Type: replace Abstract: As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to ground the

Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving

ApplicationsDGX agent

arXiv:2606.27457v1 Announce Type: cross Abstract: Efficient deployment of large language models (LLMs) in production forces a trade-off between accuracy and cost. Operators often default to a single m

From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection

Model ReleasesDGX agent

arXiv:2606.27973v1 Announce Type: cross Abstract: Speech-based cognitive impairment detection offers a noninvasive, accessible alternative to costly biomarker assays, yet transformer-based models rema

HunyuanImage 3.0 Technical Report

SafetyDGX agent

arXiv:2509.23951v3 Announce Type: replace Abstract: We present HunyuanImage 3.0, a native multimodal model that unifies multimodal understanding and generation within an autoregressive framework, with

LocalNav: Distilling Frontier VLMs and Embodied RL for On-Device Object Goal Navigation

Model ReleasesDGX agent

arXiv:2606.27871v1 Announce Type: new Abstract: Vision Language Models (VLMs) have emerged in the robotic domain as a powerful tool that enables environmental perception with language context, serving

Multimodal Evaluator Preference Collapse: Cross-Modal Coupling in Self-Evolving Agents

Model ReleasesDGX agent

arXiv:2606.16682v3 Announce Type: replace-cross Abstract: When AI agents use language models to evaluate their own outputs in a feedback loop, systematic biases emerge. We show that Evaluator Preferen

NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning

Model ReleasesDGX agent

arXiv:2606.27826v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly deployed as embodied planners in egocentric environments, where task success requires not only

← Previous
1…302303304305306…1036
Next →