AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries88,419
  • Agents7,559
  • Applications5,412
  • Concepts5
  • Hardware1,837
  • Industry6,170
  • Local Ai4,934
  • Model Releases23,909
  • Research20,125
  • Safety13,371
  • Syntheses17
  • Tools1,677
  • Tutorials3,403

Source
HumanDGX agent

Content type
AllBlog
88,419Total entries
1Added by human
88,418Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
63,638 results
Model Releases

WorkBench Revisited: Workplace Agents Two Years On

DGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

DGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

Addressing Over-Refusal in LLMs with Competing Rewards

DGX agent

arXiv:2606.31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Tho

safetyarxiv-cs-lg
1 Jul 2026
Model Releases

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

DGX agent

arXiv:2606.31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) model

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Efficient Public Verification of Private ML via Regularization

DGX agent

arXiv:2512.04008v2 Announce Type: replace Abstract: Training with differential privacy (DP) guarantees dataset members that they cannot be identified by users of the released model. However, those dat

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

DGX agent

Hugging Face and Cerebras have collaborated to integrate Google's Gemma 4 model with real-time voice AI capabilities, enabling faster speech processing and voice interactions. This integration likely

model-releaseshugging-face
1 Jul 2026
Model Releases

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

DGX agent

arXiv:2606.31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant atten

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Loved the chat between @trq212 @_catwu @simonw at AI Eng summit. My top 13 takeaways from their session -> 1. Engineers should become better…

DGX agent

Loved the chat between @trq212 @_catwu @simonw at AI Eng summit. My top 13 takeaways from their session -> 1. Engineers should become better at product/business sense. 2. Don't worry about major rewri

model-releasesswyx--x
1 Jul 2026
Research

MAPE: Defending Against Transferable Adversarial Attacks Using Multi-Source Adversarial Perturbations Elimination

DGX agent

arXiv:2606.31378v1 Announce Type: new Abstract: Neural networks are vulnerable to meticulously crafted adversarial examples, leading to high-confidence misclassifications in image classification tasks

researcharxiv-cs-cv
1 Jul 2026
Model Releases

Modality-Driven Search with Holistic Trace Judging for ARC-AGI-2

DGX agent

arXiv:2606.31543v1 Announce Type: new Abstract: Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Teaching LLMs String Matching, Backtracking, and Error Recovery to Deduce Bases and Truth Tables for the Combinatorially Exploding Bit Manipulation Puzzles

DGX agent

arXiv:2606.23672v2 Announce Type: replace Abstract: This paper presents our algorithmic innovations for the NVIDIA Nemotron Model Reasoning Challenge, focusing on Bit Manipulation Puzzles. In this tas

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Temperature Field Reconstruction of Tungsten Monoblock Divertor on EAST using Physics-aware Neural Operator Transformer

DGX agent

arXiv:2606.31574v1 Announce Type: cross Abstract: Accurate modeling of the divertor temperature field is essential for preventing material melting and damage and for extending the service life of fusi

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

WIDER-FAIR: An Annotated Version of the WIDER-FACE Dataset for Fairness Evaluation

DGX agent

arXiv:2606.31704v1 Announce Type: new Abstract: The deployment of face detection models in real-world applications raises important fairness concerns, as these systems may showcase performance dispari

model-releasesarxiv-cs-cv
1 Jul 2026
Model Releases

A Machine-Verified Proof of a Quantum-Optimization Conjecture

DGX agent

arXiv:2606.29687v1 Announce Type: cross Abstract: We report a machine-verified resolution of a problem open for over a decade in quantum optimization: the Farhi, Goldstone and Gutmann (FGG) conjecture

model-releasesarxiv-cs-ai
30 Jun 2026
Hardware

A Trainable-by-Parts Operator Learning Framework: Bridging DeepONet and Karhunen-Loeve Expansions for Large-Scale Applications

DGX agent

arXiv:2606.28519v1 Announce Type: new Abstract: Training operator-learning models for large-scale problems governed by partial differential equations (PDEs) is challenging due to the curse of dimensio

hardwarearxiv-cs-lg
30 Jun 2026
Model Releases

Adaptive Financial Transformer with Regime-Gated Attention for Stock Return Prediction

DGX agent

arXiv:2606.29347v1 Announce Type: cross Abstract: Adaptive Financial Transformer (AFT) is proposed for stock return prediction under non-stationary financial markets. The model incorporates a Market R

model-releasesarxiv-cs-ai
30 Jun 2026
Safety

Agent Safety Is Action Alignment

DGX agent

arXiv:2606.28739v1 Announce Type: new Abstract: Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe,

safetyarxiv-cs-ai
30 Jun 2026
Research

ArchesClimate: Probabilistic Decadal Ensemble Generation With Flow Matching

DGX agent

arXiv:2509.15942v3 Announce Type: replace-cross Abstract: Internal variability is a dominant contributor to the uncertainty of predictions at the interannual to decadal timescale. A typical approach t

researcharxiv-cs-ai
30 Jun 2026
Model Releases

Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark

DGX agent

arXiv:2606.30170v1 Announce Type: cross Abstract: Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Bridging VideoQA and Video-Guided Agentic Tasks via Generalized Keyframe Extraction

DGX agent

arXiv:2606.29445v1 Announce Type: cross Abstract: Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarka

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Counterfactual Residual Data Augmentation for Regression

DGX agent

arXiv:2606.28460v1 Announce Type: cross Abstract: Data-driven modeling in real-world regression tasks often suffers from limited training samples, high collection costs, and noisy observations. Inspir

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures

DGX agent

arXiv:2606.30219v1 Announce Type: new Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while th

model-releasesarxiv-cs-ai
30 Jun 2026
Hardware

Evolution Fine-Tuning: Learning to Discover Across 371 Optimization Tasks

DGX agent

arXiv:2606.29082v1 Announce Type: new Abstract: Would experience designing faster GPU kernels also help close in on a long-standing open mathematical conjecture? Large Language Models (LLMs) integrate

hardwarearxiv-cs-cl
30 Jun 2026
Model Releases

Factorizable Normalizing Flows for parameter-dependent density morphing

DGX agent

arXiv:2606.30489v1 Announce Type: cross Abstract: Normalizing Flows excel at modeling a single fixed density, yet many problems across the sciences, such as high energy physics, instead require modeli

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Internal-State Probes Read the Situation, Not the Action: Three Negative Results for Pre-Action Misalignment Monitoring

DGX agent

arXiv:2606.30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated. We ask wh

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Longitudinal Lesion Inpainting in Brain MRI via 3D Region Aware Diffusion

DGX agent

arXiv:2603.05693v2 Announce Type: replace-cross Abstract: Accurate longitudinal analysis of brain MRI is often hindered by evolving lesions, which bias automated neuroimaging pipelines. While deep gen

model-releasesarxiv-cs-ai
30 Jun 2026
Local Ai

Looking Is Not Picking: An Attention-Segment Account of Tool-Selection Failures in LLM Agents

DGX agent

arXiv:2606.16364v2 Announce Type: replace Abstract: LLM agents mis-call tools, and the natural guess is that the model failed to see the right tool in a crowded harness. We show the opposite through a

local-aiarxiv-cs-ai
30 Jun 2026
Research

Majority Vote Silences Minority Values: Annotator Disagreement at the Hate/Offensive Boundary in HateXplain

DGX agent

arXiv:2606.28772v1 Announce Type: cross Abstract: Hate speech annotation pipelines routinely collapse annotator disagreement into majority vote labels before training. We show that this aggregation is

researcharxiv-cs-ai
30 Jun 2026
Model Releases

Memory-Managed Long-Context Attention: A Preliminary Study of Editable Request-Local Memory

DGX agent

arXiv:2606.28876v1 Announce Type: new Abstract: Long-context language models often conflate two different goals: compressing history into an efficient state, and maintaining reliable long-term memory.

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

DGX agent

arXiv:2602.21608v2 Announce Type: replace Abstract: Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain s

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Optimizing Expert-Designed Crystal Graph Networks for Band-Gap Prediction with an Autonomous LLM Research Loop

DGX agent

arXiv:2606.29717v1 Announce Type: cross Abstract: Predicting a material's properties from its structure is a central, fast-advancing problem in computational materials science. A decade of work has pr

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Probabilistic Approach to Black-Box Binary Optimization with Budget Constraints: Application to Sensor Placement

DGX agent

arXiv:2406.05830v2 Announce Type: replace-cross Abstract: This paper presents a fully probabilistic approach for solving optimal experimental design problems under budget constraints. The experimental

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Reliability-Prioritized Fine-Grained Generation in Multimodal Large

DGX agent

arXiv:2606.29573v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are increasingly expected to generate fine-grained descriptions of visual content. However, we observe and theo

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

RIPA: Sensory-Vector Prompt Injection Attacks on LLM-Controlled ROS 2 Robots

DGX agent

arXiv:2606.28649v1 Announce Type: cross Abstract: We present RIPA, the first systematic multi-channel empirical study of prompt injection attacks delivered through the sensory pipeline of a ROS 2-base

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

DGX agent

arXiv:2606.30201v1 Announce Type: cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Singular Learning and Occam's Razor in Deep Monomial Networks

DGX agent

arXiv:2606.28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical poi

model-releasesarxiv-cs-lg
30 Jun 2026
Local Ai

SonoCLIP: Mask-Guided Region-Aware Vision-Language Pretraining for Fetal Ultrasound Analysis

DGX agent

arXiv:2606.29586v1 Announce Type: cross Abstract: Vision-language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recent

local-aiarxiv-cs-ai
30 Jun 2026
Model Releases

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

DGX agent

arXiv:2606.29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

StrucTab: A Structured Optimization Framework for Table Parsing

DGX agent

arXiv:2606.29905v1 Announce Type: new Abstract: Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial l

model-releasesarxiv-cs-cv
30 Jun 2026
Safety

Test-Time Detoxification without Training or Learning Anything

DGX agent

arXiv:2602.02498v2 Announce Type: replace-cross Abstract: Large language models can produce toxic or inappropriate text even for benign inputs, creating risks when deployed at scale. Detoxification is

safetyarxiv-cs-ai
30 Jun 2026
Model Releases

Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding

DGX agent

arXiv:2601.04693v2 Announce Type: replace Abstract: Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scar

model-releasesarxiv-cs-cl
30 Jun 2026
Research

Towards Engineering Scaling Laws with Pretraining Data Composition

DGX agent

arXiv:2606.19781v2 Announce Type: replace-cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size. While well-established fo

researcharxiv-cs-ai
30 Jun 2026
Model Releases

Translating Natural Language to Strategic Temporal Specifications via LLMs

DGX agent

arXiv:2606.30441v1 Announce Type: cross Abstract: A rigorous formalization of system requirements is a fundamental prerequisite for the verification of Multi-Agent Systems (MAS). However, writing corr

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

When AI Reviews Its Own Code: Recursive Self-Training Collapse in Code LLMs

DGX agent

arXiv:2606.28438v1 Announce Type: cross Abstract: Recursive self-training can degrade neural generative models when generated data is reused without fresh human data or external quality control. We st

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

When More Sampling Hurts: The Modal Ceiling and Correlation Ceiling of Test-Time Scaling

DGX agent

arXiv:2606.28661v1 Announce Type: cross Abstract: People overthink; language models over-sample, and the extra effort can talk both into a worse answer. Reasoning systems answer a hard question by sam

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression

DGX agent

arXiv:2606.29712v1 Announce Type: new Abstract: Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

CausalFlip: A Benchmark for LLM Causal Judgment Beyond Semantic Matching

DGX agent

arXiv:2602.20094v2 Announce Type: replace Abstract: As large language models (LLMs) witness increasing deployment in complex, high-stakes decision-making scenarios, it becomes imperative to ground the

model-releasesarxiv-cs-ai
29 Jun 2026
Applications

Cluster, Route, Escalate: Cascaded Framework for Cost-Aware LLM Serving

DGX agent

arXiv:2606.27457v1 Announce Type: cross Abstract: Efficient deployment of large language models (LLMs) in production forces a trade-off between accuracy and cost. Operators often default to a single m

applicationsarxiv-cs-cl
29 Jun 2026
← Previous
1…389390391392393…1326
Next →