AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,620 results
2 Jun 2026

What to Format and How: A Benchmark and Workflow Approach for Document Formatting

Model ReleasesDGX agent

arXiv:2606.01936v1 Announce Type: new Abstract: Recent advances in large language models (LLMs) have opened up new possibilities for automated document formatting. However, real-world formatting often

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Model ReleasesDGX agent

arXiv:2602.16763v2 Announce Type: replace Abstract: Artificial intelligence benchmarks are an important mechanism for measuring model progress and guiding deployment decisions. However, benchmarks qui

When and How Much to Imagine: Adaptive Test-Time Scaling with World Models for Visual Spatial Reasoning

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2602.08236v2 Announce Type: replace-cross Abstract: Despite rapid progress in MLLMs, visual spatial reasoning remains unreliable when correct answers depend on how a scene would appear under uns

When Jokes Cross the Line: Analyzing Regular Humor and Dark Humor in YouTube Shorts

Model ReleasesDGX agent

arXiv:2606.00046v1 Announce Type: cross Abstract: Video platforms such as YouTube have reshaped how users engage with entertainment and information, emphasizing brief, highly engaging content such as

When Parallelism Pays Off: Cohesion-Aware Task Partitioning for Multi-Agent Coding

Model ReleasesDGX agent

arXiv:2606.00953v1 Announce Type: new Abstract: Multi-agent Large Language Model (LLM) systems offer a way to decompose complex tasks, such as coding, through parallelization and context isolation. Ho

When Safe Skills Collide: Measuring Compositional Risk in Agent Skill Ecosystems

Model ReleasesDGX agent

arXiv:2606.00448v1 Announce Type: cross Abstract: LLM agents increasingly rely on community-contributed skills that expand an agent's operational capability set. We study a core safety problem in agen

When Single Answer Is Not Enough: Rethinking Single-Step Retrosynthesis Benchmarks for LLMs

Model ReleasesDGX agent

arXiv:2602.03554v2 Announce Type: replace-cross Abstract: Recent progress has expanded the use of large language models (LLMs) in drug discovery, including synthesis planning. However, objective evalu

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

Model ReleasesDGX agent

arXiv:2606.02060v1 Announce Type: new Abstract: Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final ans

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

Model ReleasesDGX agent

arXiv:2606.01247v1 Announce Type: new Abstract: Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has la

Which Leakage Types Matter? A Quantitative Landscape Across 2,047 Benchmark Datasets

Model ReleasesDGX agent

arXiv:2604.04199v2 Announce Type: replace Abstract: Twenty-eight within-subject counterfactual experiments across 2,047 iid tabular datasets, plus a boundary experiment on 129 temporal datasets, measu

Why Not Hyperparameter-Friendly Optimisation? A Monotonic Adaptive Norm Rescaling Approach For Long-Tailed Recognition

Model ReleasesDGX agent

arXiv:2606.02526v1 Announce Type: cross Abstract: Long-tailed recognition poses a significant challenge for deep learning. The two-stage decoupling paradigm, which separates representation learning fr

WildCat: Near-Linear Attention in Theory and Practice

Model ReleasesDGX agent

arXiv:2602.10056v2 Announce Type: replace Abstract: We introduce WildCat, a high-accuracy, low-cost approach to compressing the attention mechanism in neural networks. While attention is a staple of m

Wordle 1,808 5/6 ⬛⬛⬛⬛⬛ 🟨🟨⬛⬛⬛ ⬛🟨🟩🟨⬛ 🟩🟩🟩🟩⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

This post shows the solution path for Wordle puzzle #1,808, solved in 5 guesses, with a visual representation of each guess's results using colored tiles indicating correct letters (green), misplaced

Workflows are the biggest upgrade to Claude Code’s capabilities since skills and subagents. I dove deep into it with @sidbid to figure out b…

Model ReleasesDGX agent

Workflows are the biggest upgrade to Claude Code’s capabilities since skills and subagents. I dove deep into it with @sidbid to figure out best practices, examples and more. I’m particularly excited a

WorldCache: Accelerating World Models for Free via Heterogeneous Token Caching

Model ReleasesDGX agent

arXiv:2603.06331v2 Announce Type: replace Abstract: Diffusion-based world models have shown strong potential for unified world simulation, but the iterative denoising remains too costly for interactiv

WorldCoder-Bench: Benchmarking Physically Grounded 3D World Synthesis

Model ReleasesDGX agent

arXiv:2606.01869v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly asked not only to write static interfaces, but to construct executable interactive worlds from natural lan

WorldLens: Full-Spectrum Evaluations of Driving World Models in Real World

Model ReleasesDGX agent

arXiv:2512.10958v2 Announce Type: replace Abstract: Generative world models are reshaping embodied AI, enabling agents to synthesize realistic 4D driving environments that look convincing but often fa

Worlds Within Words: Translating Culture in Ancient Chinese Texts with Multi-Agent Coordination

Model ReleasesDGX agent

arXiv:2606.01276v1 Announce Type: new Abstract: Large language model (LLM)-based machine translation has advanced cross-cultural communication, yet it still struggles with culture-loaded words (CLWs)

WUSH: Near-Optimal Adaptive Transforms for LLM Quantization

Model ReleasesDGX agent

arXiv:2512.00956v3 Announce Type: replace Abstract: Quantizing LLM weights and activations is a standard approach for efficient deployment, but a few extreme outliers can stretch the dynamic range and

X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding

Model ReleasesDGX agent

arXiv:2606.02482v1 Announce Type: new Abstract: While video streaming understanding has made significant strides, real-world applications, such as live sports broadcasting, autonomous driving, and mul

XAI-SOH-FL: Enhancing SOH-FL with Adaptive Aggregation and Explainable AI for Intrusion Detection in Heterogeneous IoT

Model ReleasesDGX agent

arXiv:2606.00134v1 Announce Type: cross Abstract: Intrusion Detection Systems (IDS) in Internet of Things (IoT) environments face significant challenges due to data heterogeneity, lack of labeled data

You Can Learn Tokenization End-to-End with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2602.13940v2 Announce Type: replace-cross Abstract: Tokenization is a hardcoded compression step which remains in the training pipeline of Large Language Models (LLMs), despite a general trend t

Zamba2-VL Technical Report

Model ReleasesDGX agent

arXiv:2606.00390v1 Announce Type: cross Abstract: We present Zamba2-VL, a suite of vision-language models built on Zamba2, a hybrid language-model architecture combining Mamba2 state-space layers with

Zero-Shot Off-Policy Learning

Model ReleasesDGX agent

arXiv:2602.01962v2 Announce Type: replace-cross Abstract: Off-policy learning methods seek to derive an optimal policy directly from a fixed dataset of prior interactions. This objective presents sign

1 Jun 2026

3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark

Model ReleasesDGX agent

arXiv:2605.30469v1 Announce Type: cross Abstract: 3D audio and novel-view acoustic synthesis models are usually evaluated with global metrics.However, global metrics often hide where and why binaural

A Kinetic Energy Perspective of Flow Matching

Model ReleasesDGX agent

arXiv:2602.07928v2 Announce Type: replace-cross Abstract: Flow-based generative models can be viewed through a physics lens: sampling transports a particle from noise to data by integrating a learned

A Lightweight Ensemble-Based Face Image Quality Assessment Method with Correlation-Aware Loss

Model ReleasesDGX agent

arXiv:2509.10114v2 Announce Type: replace Abstract: Face image quality assessment (FIQA) plays a critical role in face recognition and verification systems, especially in uncontrolled, real-world envi

A Novel Global Context-aware Deep Neural Network for Enhanced Brain Tumor Segmentation using Magnetic Resonance Images

Model ReleasesDGX agent

arXiv:2605.30510v1 Announce Type: cross Abstract: Brain cancer's severity necessitates precise brain tumor segmentation, which is crucial for effective brain tumor diagnosis. Manual identification, bu

A Visually Impaired Assistance Benchmark for VLM-as-a-Judge Evaluation

Model ReleasesDGX agent

arXiv:2605.31351v1 Announce Type: new Abstract: AI-based Visually Impaired Assistance (VIA) remains challenging, largely due to the high cost of human evaluation. The VLM-as-a-Judge paradigm may offer

AbstainGNN: Teaching Graph Neural Networks to Abstain for Graph Classification

Model ReleasesDGX agent

arXiv:2605.30786v1 Announce Type: new Abstract: Graph classification is a core task in graph data mining with widespread real-world applications. Recent advances in graph neural networks (GNNs) have l

Adaptive NAD: Online and Self-adaptive Unsupervised Network Anomaly Detector

Model ReleasesDGX agent

arXiv:2410.22967v5 Announce Type: replace Abstract: The widespread usage of the Internet of Things (IoT) has raised the risks of cyber threats; thus, developing Anomaly Detection Systems (ADSs) that c

Aggregation Buffer: Revisiting DropEdge with a New Parameter Block

Model ReleasesDGX agent

arXiv:2505.20840v2 Announce Type: replace Abstract: We revisit DropEdge, a data augmentation technique for GNNs which randomly removes edges to expose diverse graph structures during training. While b

AMix-2: Establishing Protein as a Native Modality in Large Language Models

Model ReleasesDGX agent

arXiv:2605.30963v1 Announce Type: cross Abstract: We present AMix-2, a protein-text foundation model that establishes protein as a native modality in large language models (LLMs), unifying protein und

AMNESIA: A Large Scale Medical Unlearning Benchmark Suite with Disease-Informed Analysis

Model ReleasesDGX agent

arXiv:2605.30599v1 Announce Type: cross Abstract: Medical knowledge is continuously evolving. This creates a need to update or selectively forget information encoded in already-trained medical LLMs. M

An Odd Estimator for Shapley Values

Model ReleasesDGX agent

arXiv:2602.01399v2 Announce Type: replace-cross Abstract: The Shapley value is a ubiquitous framework for attribution in machine learning, encompassing feature importance, data valuation, and causal i

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

Model ReleasesDGX agent

arXiv:2605.30804v1 Announce Type: new Abstract: We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for Engl

Auditing LLM Benchmarks with Item Response Theory

Model ReleasesDGX agent

arXiv:2605.30504v1 Announce Type: new Abstract: LLM benchmark labels are frozen at release and silently propagated into downstream benchmarks, errors and all. We introduce an Item Response Theory-base

Auto-Discovery-Bench: Diagnosing Structured State Tracking in Oracle-Guided Discovery

Model ReleasesDGX agent

arXiv:2502.15224v2 Announce Type: replace-cross Abstract: Interactive discovery requires agents to maintain and update structured beliefs over many rounds of feedback. Before evaluating agents in nois

Automated Prediction of Postoperative Pancreatic Fistula Using Preoperative Computed Tomography

Model ReleasesDGX agent

arXiv:2605.31539v1 Announce Type: new Abstract: Postoperative pancreatic fistula (POPF) is a serious complication after pancreatic resection, increasing morbidity, hospital stay, and healthcare costs.

Automating Formal Verification with Reinforcement Learning and Recursive Inference

Model ReleasesDGX agent

arXiv:2605.30914v1 Announce Type: new Abstract: Automated formal verification remains challenging for large language models because data for proof assistants and verification-aware languages is scarce

Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence

Model ReleasesDGX agent

arXiv:2605.31484v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multi

Bandwidth Allocation with Device Partitioning for Federated Learning over Industrial IoT networks

Model ReleasesDGX agent

arXiv:2605.30892v1 Announce Type: new Abstract: We consider a federated learning (FL) system in which Industrial Internet-of-Things (IIoT) devices collaboratively train a global model over wireless ch

been asking others at Anthropic how they stay in the loop with Claude and fully understand the work being done this is one of my favorites f…

Model ReleasesDGX agent

I cannot provide an accurate summary as the title appears to be truncated and the full content is not accessible. Based on the available text, this likely discusses internal practices at Anthropic reg

Benchmarking and Enhancing Text-to-Image Models for Generating Visual Representations in Early Arithmetic Education

Model ReleasesDGX agent

arXiv:2605.31212v1 Announce Type: cross Abstract: AI systems are increasingly used to support educational content creation, yet it remains unclear whether they can generate outputs that faithfully rep

Benchmarking Uncertainty and its Disentanglement in multi-label Chest X-Ray Classification

Model ReleasesDGX agent

arXiv:2508.04457v2 Announce Type: replace-cross Abstract: Reliable uncertainty quantification is crucial for trustworthy decision-making and the deployment of AI models in medical imaging. While prior

BenHalluEval: A Multi-Task Hallucination Evaluation Framework for Large Language Models on Bengali

Model ReleasesDGX agent

arXiv:2605.31483v1 Announce Type: new Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LL

Bernini released. Unified Video generation and editing model. Built on Wan-2.2

Model ReleasesDGX agent

Bernini is a unified framework for video editing and video generation , built using Wan2.2-A14B as its renderer . The model covers complementary task families that demonstrate its capabilities as a un

Beyond Agreement: Scoring Panel-Surfaced Biomedical Entity Candidates for Curator Triage

Model ReleasesDGX agent

arXiv:2605.30826v1 Announce Type: cross Abstract: Biomedical NER is deceptively simple for modern LLMs: plausible biomedical mentions are easy to surface, but corpus-convention correctness depends on

Beyond ReLU: Bifurcation, Oversmoothing, and Topological Priors

Model ReleasesDGX agent

arXiv:2602.15634v2 Announce Type: replace Abstract: Graph Neural Networks (GNNs) learn node representations through iterative network-based message-passing. While powerful, deep GNNs suffer from overs

Beyond Static Dialogues: Benchmarking Realistic, Heterogeneous, and Evolving Long-Term Memory

Model ReleasesDGX agent

arXiv:2605.31086v1 Announce Type: new Abstract: In existing memory benchmarks for Large Language Models (LLMs), the evaluated dialogue sessions often lack long-term semantic consistency, and the under

Big day for American open models... Nemotron 3 Ultra is now the strongest US open-weight model tested, while apparently serving 300+ tok/s …

Model ReleasesDGX agent

Big day for American open models... Nemotron 3 Ultra is now the strongest US open-weight model tested, while apparently serving 300+ tok/s 🤯 Comparable large DeepSeek/Kimi models are usually 50-100 to

BilliardPhys-Bench: Benchmarking Physical Reasoning and Visual Dynamics of Multimodal LLMs

Model ReleasesDGX agent

arXiv:2605.30900v1 Announce Type: new Abstract: Current multimodal models handle static image recognition well, but intuitive physical reasoning remains a weakness. Predicting how objects will move an

Binance launches trading for 7,000+ US stocks and ETFs for non-US users, with zero commissions and fractional share purchases, as part of its 'super app' push (Jeff John Roberts/Fortune)

Model ReleasesDGX agent

Jeff John Roberts / Fortune: Binance launches trading for 7,000+ US stocks and ETFs for non-US users, with zero commissions and fractional share purchases, as part of its “super app” push — Binance, t

BlueFin: Benchmarking LLM Agents on Financial Spreadsheets

Model ReleasesDGX agent

arXiv:2605.30907v1 Announce Type: cross Abstract: We present BlueFin, a benchmark that tasks large language model (LLM) agents with synthesis, manipulation, and comprehension tasks over spreadsheet wo

BOKBO (Best of K Bad Options): Calibrated Abstention for VLA Policies

Model ReleasesDGX agent

arXiv:2605.30660v1 Announce Type: new Abstract: Test-time scaling for vision-language-action (VLA) policies, methods such as RoboMonkey, SEAL, MG-Select, and V-GPS, samples K candidate action chunks a

Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies

Model ReleasesDGX agent

arXiv:2512.19673v3 Announce Type: replace-cross Abstract: Existing reinforcement learning (RL) approaches treat large language models (LLMs) as a unified policy, overlooking their internal mechanisms.

Bounded Behavioral Indistinguishability for Black-Box LLM Distillation

Model ReleasesDGX agent

arXiv:2605.30448v1 Announce Type: cross Abstract: Black-box LLM distillation is usually evaluated as an output-matching problem: a student is considered successful when its responses are semantically

Breaking the Simplification Bottleneck in Amortized Neural Symbolic Regression

Model ReleasesDGX agent

arXiv:2602.08885v5 Announce Type: replace-cross Abstract: Symbolic regression (SR) aims to discover interpretable analytical expressions that accurately describe observed data. Amortized SR promises t

Building the infrastructure for the Intelligence Age in Michigan

Model ReleasesDGX agent

OpenAI announced plans to build significant AI infrastructure in Michigan to support the growing computational demands of advanced AI systems. The project, referred to as Stargate, represents investme

Calibrated Preference Learning: The Case of Label Ranking

Model ReleasesDGX agent

arXiv:2605.30447v1 Announce Type: cross Abstract: Calibration, the alignment of predicted probabilities with true outcome frequencies, is essential for reliable decision-making. While extensively stud

← Previous
1…186187188189190…377
Next →