AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,042
  • Agents7,457
  • Applications5,325
  • Concepts5
  • Hardware1,802
  • Industry6,143
  • Local Ai4,863
  • Model Releases23,393
  • Research19,837
  • Safety13,180
  • Syntheses17
  • Tools1,671
  • Tutorials3,349

Source
HumanDGX agent

87,042Total entries
1Added by human
87,041Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,509 results
1 Jul 2026

good post on the value of knowing whats going on in your harness 'For agents within an actual product, I have gravitated to using LangChain …

Model ReleasesDGX agent

good post on the value of knowing whats going on in your harness 'For agents within an actual product, I have gravitated to using LangChain DeepAgents' The frontier model lab harnesses are amazing - d

Histogram-constrained Image Generation

Local AiDGX agent

arXiv:2606.31683v1 Announce Type: cross Abstract: Diffusion models have emerged as a dominant paradigm in generative modeling, enabling high-fidelity sampling from complex data distributions. Despite

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding

Model ReleasesDGX agent

arXiv:2601.08758v4 Announce Type: replace-cross Abstract: Chain-of-Thought (CoT) reasoning has proven effective in enhancing large language models by encouraging step-by-step intermediate reasoning, a

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MECoBench: A Systematic Study of Multimodal Agent Collaboration in Embodied Environments

Model ReleasesDGX agent

arXiv:2606.31966v1 Announce Type: cross Abstract: Recent multimodal large language models (MLLMs) have strong potential as embodied agents, but their ability to collaborate in visually grounded enviro

Medical Image Spatial Grounding with Semantic Sampling

Model ReleasesDGX agent

arXiv:2603.14579v3 Announce Type: replace Abstract: Vision language models (VLMs) have shown significant promise in visual grounding for images as well as videos. In medical imaging research, VLMs rep

Nemotron 3 Ultra is taking off on Together AI 🚀 In just days, it's climbed to 35B tokens/day on @OpenRouter. That's the community voting wi…

Model ReleasesDGX agent

Nemotron 3 Ultra is taking off on Together AI 🚀 In just days, it's climbed to 35B tokens/day on @OpenRouter. That's the community voting with their tokens for open models that are fast, efficient, and

No Place to Hide: Benchmarking Video Hallucination with Background-Controlled Pairs

Model ReleasesDGX agent

arXiv:2606.31933v1 Announce Type: new Abstract: We introduce VidPair-Halluc, a new benchmark for evaluating video hallucination in large video models (LVMs) under rigorous and controlled conditions. U

Probing Memorization of Tabular In-Context Learning

ApplicationsDGX agent

arXiv:2606.31208v1 Announce Type: new Abstract: Large tabular models (LTMs), i.e., tabular foundation models leveraging in-context learning (ICL), achieve state-of-the-art performance on tabular tasks

Sequential RC-TGAN: Generating Relational Time Series with Spectral Envelope Loss

Model ReleasesDGX agent

arXiv:2606.31904v1 Announce Type: new Abstract: The generation of synthetic relational databases often involves modeling complex temporal dynamics, such as transaction logs or event sequences. A signi

SpikeLogBERT: Energy-Efficient Log Parsing Using Spiking Transformer Networks

Model ReleasesDGX agent

arXiv:2606.31781v1 Announce Type: cross Abstract: Log parsing is a fundamental step in automated log analysis, transforming raw system logs into structured event templates for downstream tasks such as

Tailored minimal reservoir computing: on the bidirectional connection between nonlinearities in the reservoir and in data

Model ReleasesDGX agent

arXiv:2504.17503v2 Announce Type: replace Abstract: We study how the degree of nonlinearity in the input data affects the optimal design of reservoir computers, focusing on how closely the model's non

This is exactly why we believe in customization. Quick context: Factory's original secret scanner was deterministic, so it either flagged th…

Model ReleasesDGX agent

This is exactly why we believe in customization. Quick context: Factory's original secret scanner was deterministic, so it either flagged things that weren't actually secrets (false positives) or miss

Xiaomi-GUI-0 Technical Report

Model ReleasesDGX agent

arXiv:2606.31410v1 Announce Type: new Abstract: Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions s

30 Jun 2026

AB-RAG: Adaptive Budgeted Retrieval-Augmented Generation for Reliable Question Answering

SafetyDGX agent

arXiv:2606.29090v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has become the standard way to ground large language models in external knowledge, yet most systems retrieve a fi

AerialMetric: Benchmarking and Adapting UAV Monocular Metric Depth Estimation in the Real World

Model ReleasesDGX agent

arXiv:2606.29716v1 Announce Type: new Abstract: This paper addresses the problem of monocular metric depth estimation in aerial UAV imagery. Although recent data-driven methods have achieved remarkabl

An Agentic AI Pipeline for Appliance-Level Energy Anomaly Detection and LLM-Driven Recommendations

Model ReleasesDGX agent

arXiv:2606.28467v1 Announce Type: cross Abstract: Appliance-level energy monitoring in office buildings produces noisy alerts that non-expert facility managers struggle to use. This paper proposes an

Attribution Graphs and Causal Probing for Mechanistic Discovery and Bias Repair in Multimodal Generative Learning

Model ReleasesDGX agent

arXiv:2510.12957v4 Announce Type: replace-cross Abstract: We treat the internals of generative models as mechanistic objects rather than black boxes. We introduce extbf{Attribution Graphs} (AGs), whic

Building to the Test: Coding Agents Deliver What You Check, Not What You Requested

Model ReleasesDGX agent

arXiv:2606.28430v1 Announce Type: cross Abstract: Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity proble

Claude Sonnet 5 is now available in Devin Desktop and Devin CLI. Sonnet 5 pairs frontier-level coding performance with a more affordable pri…

Model ReleasesDGX agent

Claude Sonnet 5 has been integrated into Devin Desktop and Devin CLI, offering advanced coding capabilities at a more competitive price point than previous models. This release represents an update to

Claude Sonnet 5 now available on Vercel AI Gateway

Model ReleasesDGX agent

Claude Sonnet 5 is now accessible through Vercel's AI Gateway, enabling developers to integrate Anthropic's latest language model into their applications via Vercel's unified API platform. This integr

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

Model ReleasesDGX agent

arXiv:2606.28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructi

CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

Model ReleasesDGX agent

arXiv:2606.29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading. Ye

Collective cooperation without individual fidelity in LLM agents

Model ReleasesDGX agent

arXiv:2606.30454v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as agents in simulations of social systems, yet it remains unclear when their behavior can be inter

Conversational analytics in BigQuery brings trusted agentic reasoning to everyone

Model ReleasesDGX agent

Businesses run on fast decisions, but the teams who hold the answers are often buried under a backlog of routine requests, leaving users waiting in line for insights they need now. Today, we are bring

CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates

Model ReleasesDGX agent

arXiv:2512.10342v3 Announce Type: replace Abstract: Vision Language Models (VLMs) have shown promising planning capabilities, yet their success remains confined to the text domain, leaving visual deci

CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents

Model ReleasesDGX agent

arXiv:2511.02734v3 Announce Type: replace Abstract: Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability

Cross-Temporal Sinhala OCR: Page-Level Adaptation and Diachronic Analysis

Model ReleasesDGX agent

arXiv:2606.29378v1 Announce Type: new Abstract: Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world data

Customized Generative AI Agent for Transportation Engineering Practice: A Development and Continued Pre-training Guideline

Model ReleasesDGX agent

arXiv:2606.29014v1 Announce Type: new Abstract: Recent advancements in generative artificial intelligence (AI) and large language models (LLMs) have shown significant promise in automating complex rea

Database Context Compression for Text-to-SQL on Real-World Large Databases

Model ReleasesDGX agent

arXiv:2606.28601v1 Announce Type: cross Abstract: Recent progress in Text-to-SQL has been driven by stronger language models and prompting strategies, yet performance on real enterprise benchmarks suc

Detecting Clinical Hallucinations in LVLMs via Counterfactual Visual Grounding Uncertainty

Model ReleasesDGX agent

arXiv:2606.28520v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) are increasingly used for clinical image understanding, yet they remain vulnerable to hallucinations--producing t

DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects

Model ReleasesDGX agent

arXiv:2604.05318v2 Announce Type: replace Abstract: Harmful content detectors, particularly disinformation classifiers, are predominantly developed and evaluated on Standard American English (SAE), le

DistilledGemma: Balanced Efficiency-Accuracy for Person-Place Relation Extraction from Multilingual Historical Articles

Model ReleasesDGX agent

arXiv:2606.29130v1 Announce Type: new Abstract: We present DistilledGemma, an efficient and accurate system for the HIPE-2026 shared task on person-place relation extraction from multilingual historic

DLR: Zero-Inference-Cost Latent Residuals for Low-Rank Pre-Training

Model ReleasesDGX agent

arXiv:2606.28932v1 Announce Type: cross Abstract: Large language models have driven recent progress in language and multimodal AI, yet pre-training them at scale is prohibitively expensive. Low-rank p

DRIFT: Difficulty Routing Self-DIstillation with Rhythm-Gated Exploration and Success BuFfer Training

SafetyDGX agent

arXiv:2606.30345v1 Announce Type: cross Abstract: Enabling large language models to achieve stable self-improvement without external expert supervision remains a central challenge in complex reasoning

Dynamic Parsing and Updating Natural Language Specification using VLMs for Robust Vision-Language Tracking

Model ReleasesDGX agent

arXiv:2606.29357v1 Announce Type: cross Abstract: Vision-language tracking guided by natural language specifications leverages high-level semantic cues of target objects to substantially boost trackin

Dynamo: Dynamic Skill-Tool Evolution for Vision-Language Agents

Model ReleasesDGX agent

arXiv:2606.30185v1 Announce Type: new Abstract: Improving vision-language models (VLMs) on visual reasoning typically requires retraining or hand-designed prompts and tools. We present Dynamo, a train

EVLA: An Electro-Aware Multimodal Assistant for Physically-Grounded Driving Reasoning and Control

Model ReleasesDGX agent

arXiv:2606.28938v1 Announce Type: new Abstract: Modern vision-language models (VLMs) for driving assistants typically treat vehicle dynamics as a black box, resulting in decisions that lack awareness

Experience Augmented Policy Optimization for LLM Reasoning

Model ReleasesDGX agent

arXiv:2606.30420v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) is a powerful paradigm for improving the reasoning capabilities of large language models (LLMs). H

Few-Step Boltzmann Generators via Scalable Likelihood Flow Maps

ResearchDGX agent

arXiv:2606.29110v1 Announce Type: new Abstract: Recent progress in flow-based generative modeling has led to models that output high-quality samples while using only a small number of function evaluat

Generalization Analysis of Transformers in Distribution Regression

Model ReleasesDGX agent

arXiv:2606.29256v1 Announce Type: cross Abstract: In recent years, models based on the Transformer architecture have seen widespread applications and have become one of the core tools in the field of

Heterogeneous Tactile Transformer

Model ReleasesDGX agent

arXiv:2606.29948v1 Announce Type: new Abstract: Tactile sensors are inherently heterogeneous: a model trained on one sensor cannot be directly used on another, which limits learning contact-rich manip

How Far Can You Get Without a GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation

Model ReleasesDGX agent

arXiv:2606.29809v1 Announce Type: cross Abstract: Hallucination detection has become a pressing requirement for trustworthy AI deployment at scale. The most accurate detection methods depend on GPU-in

How much of an LLM-generated clinical corpus is actually new? A production-scale measurement of content redundancy for provenance classification

Model ReleasesDGX agent

arXiv:2606.29605v1 Announce Type: new Abstract: Clinical machine learning increasingly relies on training corpora generated by large language models (LLMs) rather than annotated by clinicians, and suc

I-BBS: Coordinate-Free Inference of Latent Sub-Manifolds Using Random Distance Matrix Theory

Model ReleasesDGX agent

arXiv:2606.29675v1 Announce Type: new Abstract: Bogomolny, Bohigas and Schmit (BBS) found that the spectrum of the pairwise distance matrix on N points sampled from a smooth d-dimensional manifold enc

Intermediate Text Representation Guided Text-to-Image Generation for Enhancing One-and-Only Alignment

Model ReleasesDGX agent

arXiv:2606.30262v1 Announce Type: new Abstract: Text-to-image (T2I) diffusion models often fail to faithfully render explicit textual descriptions, instead defaulting to strongly learned visual priors

Introducing GeneBench-Pro

Model ReleasesDGX agent

GeneBench-Pro is a comprehensive benchmarking tool or dataset introduced by OpenAI designed to evaluate AI model performance on genomic and gene-related tasks. It likely provides standardized metrics

LoGSAM: Parameter-Efficient Cross-Modal Grounding for MRI Segmentation

Model ReleasesDGX agent

arXiv:2603.17576v3 Announce Type: replace Abstract: Precise localization and delineation of brain tumors using magnetic resonance imaging (MRI) are essential for planning therapy and guiding surgical

LUMEN: Cost-Transparent Multi-Agent Pipeline for Automated Systematic Review and Meta-Analysis

Model ReleasesDGX agent

arXiv:2606.28362v1 Announce Type: cross Abstract: Systematic reviews and meta-analyses (SR/MA) remain the gold standard for evidence synthesis, yet completing one typically requires 67 weeks and subst

Multimodal Graph RAG for Long-range Visually Rich Document Understanding

Model ReleasesDGX agent

arXiv:2606.28780v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are widely applied to visual document understanding. However, comprehending long documents remains an issue b

Multimodal Mathematical Reasoning with Diverse Solving Perspective

Model ReleasesDGX agent

arXiv:2507.02804v2 Announce Type: replace Abstract: Recent progress in large-scale reinforcement learning (RL) has notably enhanced the reasoning capabilities of large language models (LLMs), especial

Nano Banana 2 Lite

Model ReleasesDGX agent

Nano Banana 2 Lite Also known as Gemini 3.1 Flash Lite Image (gemini-3.1-flash-lite-image in their API), this is the 'fastest and cheapest Gemini image model, engineered for velocity and scale'. I use

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

Model ReleasesDGX agent

arXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produc

OP3DSG: Open-Vocabulary Part-Aware 3D Scene Graph Generation for Real-World Environments

Model ReleasesDGX agent

arXiv:2606.29786v1 Announce Type: new Abstract: 3D scene graphs (3DSGs) provide a compact and structured abstraction of 3D environments. Although advances in foundation models have enabled open-vocabu

Perspectives on Latent Factor Indeterminacy and its Implications for Data Representation

ResearchDGX agent

arXiv:2606.28854v1 Announce Type: cross Abstract: The common factor analytic model is related to Helmholtz and Boltzmann machines, can be conceived as a linear autoencoder, or can be thought of as a s

REAR: Test-time Preference Realignment through Reward Decomposition

SafetyDGX agent

arXiv:2606.30339v1 Announce Type: new Abstract: Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to

REPAIR-Bench: A Benchmark for Robot Error Perception And Interaction Recovery

Model ReleasesDGX agent

arXiv:2606.29937v1 Announce Type: new Abstract: Understanding how users perceive and respond to robot failures is essential for building robust and trustworthy robot systems. Prior work, however, (i)

Reported Confidence in LLMs Tracks Commitment More Than Correctness

Model ReleasesDGX agent

arXiv:2606.29490v1 Announce Type: cross Abstract: Confidence is an estimate of the probability that a chosen answer is correct. Verbal confidence reports are widely used as uncertainty measures in lar

Reward-Free Code Alignment from Pretrained or Fine-Tuned LLM: Unpacking the Trade-offs for Code Generation

Model ReleasesDGX agent

arXiv:2606.28998v1 Announce Type: cross Abstract: Large Language Model (LLM) alignment trains an LLM using preference data to produce outputs that better meet established quality standards. While LLM

SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing

Model ReleasesDGX agent

arXiv:2606.29887v1 Announce Type: new Abstract: In real-world applications, guardrails are often expected to identify unsafe user-model interactions according to application-specific safety policies,

Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation

Model ReleasesDGX agent

arXiv:2606.30244v1 Announce Type: new Abstract: Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression

← Previous
1…338339340341342…1042
Next →