AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,510
  • Agents7,405
  • Applications5,305
  • Concepts5
  • Hardware1,789
  • Industry6,120
  • Local Ai4,835
  • Model Releases23,219
  • Research19,716
  • Safety13,102
  • Syntheses17
  • Tools1,670
  • Tutorials3,327

Source
HumanDGX agent

86,510Total entries
1Added by human
86,509Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,082 results
2 Jun 2026

Logit Distillation on Manifolds: Mapping by Learning

ResearchDGX agent

arXiv:2606.00771v1 Announce Type: cross Abstract: A simple way to improve the performance of almost any machine learning model is not to train a single but several models with diverse algorithms which

MCP-Persona: Benchmarking LLM Agents on Real-World Personal Applications via Environment Simulation

Model ReleasesDGX agent

arXiv:2606.02470v1 Announce Type: new Abstract: The Model Context Protocol (MCP) has emerged as a transformative standard for connecting large language models (LLMs) with external data sources and too

PaintBench: Deterministic Evaluation of Precise Visual Editing

Model ReleasesDGX agent

arXiv:2606.00188v1 Announce Type: cross Abstract: While current multimodal models are proficient at open-ended visual editing, executing precise single-answer edits remains an important obstacle. To p

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

PEACE: A Planner-Executor Agent with Constraint Enforcement for UAVs

Local AiDGX agent

arXiv:2606.00104v1 Announce Type: cross Abstract: Foundation models are increasingly used to drive autonomous systems, yet existing approaches either keep the model in a tight control loop, raising la

ProductWebGen: Benchmarking Multimodal Product Webpage Generation

Model ReleasesDGX agent

arXiv:2606.01022v1 Announce Type: cross Abstract: Crafting a product display webpage from a source product image, along with layout and visual content instructions, holds significant practical value f

ProjQ: Project-and-Quantize for Adapter-Aware LLM Compression

Model ReleasesDGX agent

arXiv:2606.00494v1 Announce Type: new Abstract: Post-Training Quantization (PTQ) and Low-Rank Adaptation (LoRA) constitute the standard pipeline for efficient Large Language Model (LLM) deployment. Ho

Prototypicality Bias Reveals Blindspots in Multimodal Evaluation Metrics

Model ReleasesDGX agent

arXiv:2601.04946v3 Announce Type: replace-cross Abstract: Automatic metrics are widely used to evaluate text-to-image models, often replacing human judgment in benchmarking, model selection, and large

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

Model ReleasesDGX agent

arXiv:2606.01216v1 Announce Type: new Abstract: The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling

Structure Enables Effective Self-Localization of Errors in LLMs

Local AiDGX agent

arXiv:2602.02416v2 Announce Type: replace Abstract: Self-correction in language models remains elusive. In this work, we explore whether language models can explicitly localize errors in incorrect rea

Subliminal Learning Is Steering Vector Distillation

ResearchDGX agent

arXiv:2606.00995v1 Announce Type: new Abstract: Subliminal learning refers to a student language model acquiring a teacher's traits (e.g. a system-prompted preference for owls) when fine-tuned on the

The Right Inference Strategy Is All You Need: Nearly Training-Free Domain-Wise Inference for EgoCross Challenge

ResearchDGX agent

arXiv:2606.00829v1 Announce Type: new Abstract: EgoCross evaluates multimodal large language models on egocentric video question answering under substantial domain shift, where test videos come from s

Toward accurate RUL and SoH estimation using reinforced graph-based physics-informed neural networks enhanced with dynamic weights

Model ReleasesDGX agent

arXiv:2507.09766v2 Announce Type: replace-cross Abstract: Accurate estimation of Remaining Useful Life (RUL) and State of Health (SoH) is essential for reliable Prognostics and Health Management (PHM)

TravelEval: A Comprehensive Benchmarking Framework for Evaluating LLM-Powered Travel Planning Agents

Model ReleasesDGX agent

arXiv:2606.01046v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has significantly improved travel planning applications, yet evaluating such models is limited by existi

TukaBench: A Culturally Grounded Jailbreak Benchmark for African Languages

Model ReleasesDGX agent

arXiv:2606.01322v1 Announce Type: cross Abstract: Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, c

Variational Learning for Insertion-based Generation

ResearchDGX agent

arXiv:2606.02133v1 Announce Type: cross Abstract: Non-monotonic sequence generation methods, such as masked diffusion models, provide a flexible alternative to left-to-right autoregressive modeling by

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

ResearchDGX agent

arXiv:2606.02378v1 Announce Type: cross Abstract: We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (de

1 Jun 2026

3DAE: Binaural Quality Assessment for Audio Novel View Synthesis with Spatial Maps and Benchmark

Model ReleasesDGX agent

arXiv:2605.30469v1 Announce Type: cross Abstract: 3D audio and novel-view acoustic synthesis models are usually evaluated with global metrics.However, global metrics often hide where and why binaural

Anchoring LLM Gender Bias to Human Baselines: A Cross-Lingual Audit

Model ReleasesDGX agent

arXiv:2605.30804v1 Announce Type: new Abstract: We audit six large language models (LLMs) for gender stereotyping across English, Korean, Chinese, and Japanese. Three were developed primarily for Engl

Automating Formal Verification with Reinforcement Learning and Recursive Inference

Model ReleasesDGX agent

arXiv:2605.30914v1 Announce Type: new Abstract: Automated formal verification remains challenging for large language models because data for proof assistants and verification-aware languages is scarce

Can LLM Teams Play What? Where? When?

Model ReleasesDGX agent

arXiv:2605.30459v1 Announce Type: new Abstract: Large language models (LLMs) remain limited on tasks requiring indirect reasoning, cultural knowledge, and coordinated hypothesis testing. We investigat

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution

Model ReleasesDGX agent

arXiv:2605.30802v1 Announce Type: cross Abstract: Prediction markets aggregate collective intelligence to forecast uncertain events, but their utility depends on reliable outcome resolution. Existing

Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding

Model ReleasesDGX agent

arXiv:2511.19923v2 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reas

DTG-Restore: Training-Free Diffusion Refinement for Generative Video Super-Resolution

Model ReleasesDGX agent

arXiv:2605.30431v1 Announce Type: new Abstract: Recent progress in video diffusion models has enabled remarkable generative fidelity, yet leveraging these priors for restoration remains limited by the

FAM-Bench: A Multimodal Benchmark for Condition-Aware Food-as-Medicine Reasoning

Model ReleasesDGX agent

arXiv:2605.31410v1 Announce Type: new Abstract: Food-as-Medicine requires models to reason beyond what a dish is or what nutrition it contains: they must decide whether a concrete food choice is appro

Forecasting with Hyper-Trees

SafetyDGX agent

arXiv:2405.07836v5 Announce Type: replace Abstract: We introduce Hyper-Trees as a novel framework for modeling time series data using gradient boosted trees. Unlike conventional tree-based approaches

Generative Models and Statistical Validation

ResearchDGX agent

arXiv:2605.30453v1 Announce Type: cross Abstract: Generative machine learning has become an essential tool in theoretical and experimental physics, especially in the context of fast surrogates and den

Graph Machine Learning in the Era of Large Language Models (LLMs)

ResearchDGX agent

arXiv:2404.14928v3 Announce Type: replace-cross Abstract: Graphs play an important role in representing complex relationships in various domains like social networks, knowledge graphs, and molecular d

Microsoft to unveil new AI models and Windows improvements at Build

IndustryDGX agent

Microsoft is heading to San Francisco this week in a bid to win back developers at its Build conference. I've been attending Build since the days when Microsoft called it the Professional Developers C

Modeling Covariate Transition for Efficient Estimation of Longitudinal Treatment Effects in Randomized Experiments

ResearchDGX agent

arXiv:2605.31443v1 Announce Type: cross Abstract: We present a regression-adjustment framework designed for the estimation of longitudinal treatment effects in randomized experiments under static regi

So much great work lately from Nvidia, the 'King of American Open-source AI'! - Crossed 1,000 total public repositories on @huggingface (820…

Model ReleasesDGX agent

So much great work lately from Nvidia, the 'King of American Open-source AI'! - Crossed 1,000 total public repositories on @huggingface (820 models, 249 datasets & 57 spaces) & almost 60,000 followers

Target-Side Paraphrase Augmentation for Sign Language Translation with Large Language Models

ResearchDGX agent

arXiv:2605.31393v1 Announce Type: cross Abstract: Sign language translation (SLT) remains constrained by limited paired sign-video/text corpora and heavy-tailed target vocabularies. We study target-si

The fully-managed Remote MCP Server for AlloyDB is now Generally Available

Model ReleasesDGX agent

AI agents possess incredible reasoning capabilities and can perform increasingly complex actions. But the reliability of agentic outcomes depends entirely on the quality of the context they can access

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

Model ReleasesDGX agent

arXiv:2602.07864v2 Announce Type: replace Abstract: Spatial intelligence is crucial for vision--language models (VLMs), yet many scene-centric benchmarks evaluate unconstrained environments where a si

This pod was an incredible gift to the community: not only our first pod about @xAI, but Ethan really indulged on all our questions on how t…

HardwareDGX agent

This pod was an incredible gift to the community: not only our first pod about @xAI, but Ethan really indulged on all our questions on how to train a SOTA Videogen world model, including specific area

What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness

Model ReleasesDGX agent

arXiv:2605.30911v1 Announce Type: cross Abstract: Hallucination remains one of the key challenges undermining the reliability of Large Vision-Language Models (LVLMs). But what makes an LVLM hallucinat

XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks

Model ReleasesDGX agent

arXiv:2605.30788v1 Announce Type: cross Abstract: We introduce a set of synthetic algorithmic tasks to detect cross-lingual gaps in the abilities of large language models. Our benchmark is commensurat

29 May 2026

11 demos of Gemini Omni and Gemini 3.5 in action

Model ReleasesDGX agent

Google unveiled Gemini 3.5 Flash, the latest model combining frontier intelligence with action, and Gemini Omni, a new model that can create anything from any input starting with video. With Omni, use

Accommodation Goes Both Ways: Studying Linguistic Convergence Between Humans and Language Models

ApplicationsDGX agent

arXiv:2605.29278v1 Announce Type: new Abstract: As LLMs become increasingly integrated into daily life, understanding how their presence will shape human linguistic behavior is an open question. We pr

Architecture-Sensitive Supervised Fine-Tuning for Screen-Conditioned Action Prediction: A PiSAR Benchmark

Model ReleasesDGX agent

arXiv:2605.29400v1 Announce Type: new Abstract: We benchmark three supervised fine-tuned models against frontier zero-shot baselines on a 661-row held-out slice of PiSAR (Persona, intent, Screen, Acti

Benchmarking Single-Factor Physical Video-to-Audio Generation

Model ReleasesDGX agent

arXiv:2605.30339v1 Announce Type: new Abstract: Generative video-to-audio (V2A) models produce highly plausible soundtracks, but it remains unclear whether they capture the underlying physical process

Beyond Trajectory Rewards: Step-level Credit Assignment for Agentic Search via Graph Modeling

SafetyDGX agent

arXiv:2605.29697v1 Announce Type: new Abstract: In Agentic Search, trajectory-level outcome rewards fail to quantify the behavioral contributions of individual steps, while existing step-level reward

Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori Estimation

ResearchDGX agent

arXiv:2605.30269v1 Announce Type: new Abstract: Over the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individ

Certified Causal Defense with Generalizable Robustness

Model ReleasesDGX agent

arXiv:2408.15451v3 Announce Type: replace Abstract: While machine learning models have proven effective across various scenarios, it is widely acknowledged that many models are vulnerable to adversari

Cloud CISO Perspectives: How to build an AI-ready security program for the public sector

Model ReleasesDGX agent

Welcome to the second Cloud CISO Perspectives for May 2026. Today, Usman Chaudhary, Field CISO, Google Public Sector, offers a guide for CISOs protecting government agencies and critical infrastructur

DFlash: Block Diffusion for Flash Speculative Decoding

HardwareDGX agent

arXiv:2602.06036v2 Announce Type: replace Abstract: Autoregressive large language models (LLMs) deliver strong performance but require inherently sequential decoding, leading to high inference latency

EarthShift: a benchmark for measuring robustness to real-world distribution shifts in Earth observation

Model ReleasesDGX agent

arXiv:2605.29330v1 Announce Type: new Abstract: Current Earth observation benchmarks focus on measuring performance on diverse tasks and applications, typically measuring generalization in-distributio

FinGuard: Detecting Financial Regulatory Non-Compliance in LLM Interactions

Model ReleasesDGX agent

arXiv:2605.29427v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly deployed in financial services, a single non-compliant interaction can expose institutions to regulator

GPIC: A Giant Permissive Image Corpus for Visual Generation

Model ReleasesDGX agent

arXiv:2605.30341v1 Announce Type: cross Abstract: Studying scalable methods for visual generative modeling requires large, accessible, and stable datasets. We introduce GPIC, a Giant Permissive Image

HoliTok:A Coutinuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding

ResearchDGX agent

arXiv:2605.29948v1 Announce Type: cross Abstract: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality wavef

K-FinHallu: A Hallucination Detection Benchmark for Multi-Turn RAG in Korean Finance

Model ReleasesDGX agent

arXiv:2605.29523v1 Announce Type: new Abstract: Large Language Models (LLMs) have advanced financial automation through Retrieval-Augmented Generation (RAG), yet hallucinations remain a critical barri

KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency Tradeoffs

ResearchDGX agent

arXiv:2605.29259v1 Announce Type: cross Abstract: Given the wide range of deployment targets, flexible model selection is essential for optimizing performance within a given compute budget. Recent wor

Less is Enough: Synthesizing Diverse Data in LLM Feature Space with Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2602.10388v3 Announce Type: replace-cross Abstract: The diversity of post-training data is critical for effective downstream performance in large language models (LLMs). Many existing approaches

MiAD: Mirage Atom Diffusion for De Novo Crystal Generation

ResearchDGX agent

arXiv:2511.14426v2 Announce Type: replace-cross Abstract: In recent years, diffusion-based models have demonstrated exceptional performance in searching for simultaneously stable, unique, and novel (S

Mind Your Tone: Does Tone Alter LLM Performance?

Model ReleasesDGX agent

arXiv:2605.29027v1 Announce Type: new Abstract: The use of Large Language Models (LLMs) is proliferating, yet their performance is observed to vary based on prompting styles and tones. In this study,

Neural Scaling Laws for Jet Generation

ResearchDGX agent

arXiv:2605.28940v1 Announce Type: cross Abstract: Recently observed empirical scaling laws describe the performance of foundation-type models as three independent key quantities -- dataset size, compu

Reasoning with Sampling: Cutting at Decision Points

Local AiDGX agent

arXiv:2605.30327v1 Announce Type: cross Abstract: Frontier reasoning models are produced by posttraining base language models with reinforcement learning. Recent work has challenged this by showing th

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

Model ReleasesDGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

Representation Unlearning: Forgetting through Information Compression

Model ReleasesDGX agent

arXiv:2601.21564v2 Announce Type: replace Abstract: Machine unlearning seeks to remove the influence of specific training data from a model, a need driven by privacy regulations and robustness concern

Scaling Laws for Agent Harnesses via Effective Feedback Compute

Model ReleasesDGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

SERC: LDPC-Inspired Semantic Error Correction for Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.28837v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have demonstrated remarkable capabilities, their reliability is significantly compromised by hallucinations. Existi

← Previous
1…280281282283284…1035
Next →