AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
5 Jun 2026

SoCRATES: Towards Reliable Automated Evaluation of Proactive LLM Mediation across Domains and Socio-cognitive Variations

Model ReleasesDGX agent

arXiv:2606.05563v1 Announce Type: cross Abstract: Evaluating LLM mediators remains challenging, as mediation unfolds as a real-time trajectory shaped by disputants' shifting emotions, intentions, and

Sources say xAI used Claude models for distillation and training, including using personal accounts and the intermediary service Blackbox AI after being cut off (Grace Kay/The Information)

Model ReleasesDGX agent

Grace Kay / The Information: Sources say xAI used Claude models for distillation and training, including using personal accounts and the intermediary service Blackbox AI after being cut off — SpaceX's

Stability vs. Manipulability: Evaluating Robustness Under Post-Decision Interaction in LLM Judges


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model ReleasesDGX agent

arXiv:2606.05384v1 Announce Type: cross Abstract: LLM-as-judge evaluation is widely used in benchmarking pipelines, where model outputs are compared and ranked using automated evaluators. These pipeli

Statistically Reliable LLM-Based Ranking Evaluation via Prediction-Powered Inference

Model ReleasesDGX agent

arXiv:2606.05308v1 Announce Type: cross Abstract: With PRECISE, we extended Prediction-Powered Inference to produce bias-corrected estimates of ranking evaluation metrics by combining a small human-la

StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset

Model ReleasesDGX agent

arXiv:2606.06338v1 Announce Type: new Abstract: Video question answering (VideoQA) aims to answer questions about given videos. While existing approaches excel on factoid VideoQA, they struggle with d

SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents

Model ReleasesDGX agent

arXiv:2606.05761v1 Announce Type: cross Abstract: Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they

TARPO: Token-Wise Latent-Explicit Reasoning via Action-Routing Policy Optimization

Model ReleasesDGX agent

arXiv:2606.05859v1 Announce Type: new Abstract: Latent reasoning has emerged as a promising alternative to discrete Chain-of-Thought (CoT) in large language models (LLMs), enabling more expressive rea

Ten Headache Specialists versus Artificial Intelligence for Clinical Literature Summarization: A Critical Evaluation and Comparison

Model ReleasesDGX agent

arXiv:2606.05436v1 Announce Type: cross Abstract: Summarizing the latest medical literature to guide clinical decision-making is essential for evidence-based medicine and high-quality patient care. Ye

TensorBench: Benchmarking Coding Agents on a Compiler-Based Tensor Framework

Model ReleasesDGX agent

arXiv:2606.05570v1 Announce Type: new Abstract: Repository-level coding benchmarks face a trade-off between task difficulty and evaluation reliability: tasks that challenge frontier models often invol

TextWand: A Unified Framework for Scene Text Editing

Model ReleasesDGX agent

arXiv:2606.05730v1 Announce Type: new Abstract: We propose TextWand, a general-purpose framework that unifies scene text removal, generation, and replacement into a single model. By decomposing comple

The Granularity Gap: A Multi-Dimensional Longitudinal Audit of Sycophancy in Gemini Models

Model ReleasesDGX agent

arXiv:2606.05183v1 Announce Type: new Abstract: Large language models are increasingly deployed as high-stakes advisors, yet standard alignment benchmarks treat sycophancy as a binary failure mode. We

The latest AI news we announced in May 2026

Model ReleasesDGX agent

Google's May 2026 AI updates center on the new 'agentic' era, featuring the Gemini 3.5 model and Gemini Omni for advanced reasoning and creation. Gemini Omni is a new model that can create anything fr

The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?

Model ReleasesDGX agent

arXiv:2504.10020v4 Announce Type: replace Abstract: Contrastive decoding strategies are widely used to reduce object hallucinations in multimodal large language models (MLLMs). These methods work by c

The NVIDIA Nemotron Coalition continues to grow. We're excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And…

Model ReleasesDGX agent

The NVIDIA Nemotron Coalition continues to grow. We're excited to welcome new members: @hcompany_ai, @NousResearch, and @PrimeIntellect. And a big thank you to our existing members: @bfl_ai, @cursor_a

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators

Model ReleasesDGX agent

arXiv:2606.06476v1 Announce Type: new Abstract: While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the

Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents ca…

Model ReleasesDGX agent

Token costs are why there will be no saas apocalypse / good dev tools are cached intelligence for agents! The popular theory goes: agents can write code, so they'll just rebuild every tool from scratc

TopoPult-SSL: Gland-Mask-Free Cross-Device Meibomian Gland Segmentation via Self-Distilled Weak Clinical Priors

Model ReleasesDGX agent

arXiv:2606.05347v1 Announce Type: new Abstract: Every new clinical imaging device creates a domain shift where dense gland masks are expensive yet cheap clinical signals -- eyelid outlines, Pult grade

Towards Accurate Heart Rate Measurement from Ultra-Short Video Clips via Periodicity-Guided rPPG Estimation and Signal Reconstruction

Model ReleasesDGX agent

arXiv:2506.22078v2 Announce Type: replace Abstract: Many remote Heart Rate (HR) measurement methods focus on estimating remote photoplethysmography (rPPG) signals from video clips lasting around 10 se

Towards One-to-Many Temporal Grounding

Model ReleasesDGX agent

arXiv:2606.06294v1 Announce Type: new Abstract: Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retriev

UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning

Model ReleasesDGX agent

arXiv:2606.05576v1 Announce Type: new Abstract: Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images -

Unifying Dataset Pruning and Distillation for Efficient Large-scale Compression

Model ReleasesDGX agent

arXiv:2502.06434v2 Announce Type: replace Abstract: Dataset pruning (DP) and dataset distillation (DD) fundamentally differ in their outputs: DP selects original image subsets, while DD generates synt

UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching

Model ReleasesDGX agent

arXiv:2606.05399v1 Announce Type: new Abstract: Existing feed-forward networks excel at predicting a single set of physical properties from visual appearance, but this point-estimate paradigm fundamen

Unlocking dependable responses with Gemini Enterprise Agent Platform’s Agentic RAG

Model ReleasesDGX agent

Google's RAG Engine securely connects private enterprise data to LLMs to improve answer accuracy and reduce hallucinations , making it a key component of the Gemini Enterprise Agent Platform for build

Using Large Language Models to Support High Volume Application Review for an Undergraduate Research Program

Model ReleasesDGX agent

arXiv:2606.05564v1 Announce Type: new Abstract: Undergraduate research programs such as the Summer Undergraduate Research Fellowship (SURF) at Purdue University receive thousands of applications every

V2V-Bench: A Comprehensive Benchmark for Video-to-Video Generation Evaluation

Model ReleasesDGX agent

arXiv:2606.05665v1 Announce Type: new Abstract: Video-to-video (V2V) generation is difficult to evaluate because outputs must both follow editing instructions and preserve frame-level correspondence w

Video-Rate Streaming Stylization on a Vision-Aware MLLM-Conditioned Edit Diffusion: Asymmetric Batched Inference on a Distilled UNet + MLLM Text Encoder

Model ReleasesDGX agent

arXiv:2606.05981v1 Announce Type: new Abstract: Aggressive distillation of the diffusion U-Net inverts the per-frame bottleneck of real-time text-to-image pipelines: once the denoiser is a 4-step or 1

VideoKR: Towards Knowledge- and Reasoning-Intensive Video Understanding

Model ReleasesDGX agent

arXiv:2606.05259v1 Announce Type: new Abstract: We introduce VideoKR, the first large-scale training corpus specifically designed to strengthen knowledge- and reasoning-intensive video understanding.

VTI-CoT: Visual-Textual Interleaved Chain of Thought for Video Reasoning

Model ReleasesDGX agent

arXiv:2606.05736v1 Announce Type: new Abstract: Video reasoning aims to understand complex temporal events and causal relationships within videos. Recently, Chain-of-Thought (CoT) has been introduced

VZCrash: A Large-Scale IMU Dataset of Ego-Vehicle Crashes

Model ReleasesDGX agent

arXiv:2606.06074v1 Announce Type: new Abstract: We introduce VZCrash, the largest publicly available dataset of real-world vehicle collision data featuring Inertial Measurement Unit (IMU) telemetry. T

Waypoints Matter: A Systematic Study for Sampling-Based Trajectory Planning

Model ReleasesDGX agent

arXiv:2606.06366v1 Announce Type: new Abstract: Real-time autonomous driving commonly relies on sampling-based trajectory planners that link candidate trajectories to target waypoints along the road c

We doubled Claude Cowork usage limits for the next month. This applies to your 5-hr rate limits. If you’ve been saving up a big messy projec…

Model ReleasesDGX agent

We doubled Claude Cowork usage limits for the next month. This applies to your 5-hr rate limits. If you’ve been saving up a big messy project, now’s the time. We've doubled usage limits in Claude Cowo

We've made a breakthrough in self-evolving AI scientists moving from 'search' to 'principled discovery': Scientific discovery requires that …

Model ReleasesDGX agent

We've made a breakthrough in self-evolving AI scientists moving from 'search' to 'principled discovery': Scientific discovery requires that the search space itself changes, and an AI scientist must pe

Wordle 1,811 5/6 ⬛⬛🟨⬛⬛ 🟨⬛⬛⬛🟨 ⬛🟩🟨🟨⬛ 🟩🟩🟩🟩⬛ 🟩🟩🟩🟩🟩

Model ReleasesDGX agent

Anthropic's official X account shared a Wordle game result (puzzle #1,811) solved in 5 attempts, displaying the color-coded tile progression from each guess. The post documents the company's engagemen

Would you still call this Dax? Novel Visual References in VLMs and Humans

Model ReleasesDGX agent

arXiv:2606.05409v1 Announce Type: cross Abstract: Vision-language models (VLMs), like human learners, are frequently exposed to new visual concepts, but how they map novel visual references to languag

Your AI chatbot is only as good as the data behind it. This n8n template from our friends at @apify shows you how to wire up a RAG pipeline …

Model ReleasesDGX agent

Your AI chatbot is only as good as the data behind it. This n8n template from our friends at @apify shows you how to wire up a RAG pipeline using Apify + Pinecone + Gemini so your chatbot can answer q

YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition

Model ReleasesDGX agent

arXiv:2606.05868v1 Announce Type: new Abstract: Large language models (LLMs) drive significant financial innovations, yet their high-concurrency deployment is severely bottlenecked by KV cache memory

4 Jun 2026

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?

Model ReleasesDGX agent

arXiv:2505.19293v2 Announce Type: replace-cross Abstract: Long-context capability is considered one of the most important abilities of LLMs, as a truly long context-capable LLM enables users to effort

4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filte…

Model ReleasesDGX agent

4/5 Still came in ~$0.30 under Claude Code’s spend at a similar score. So we added a lightweight Test Agent that writes repo tests and filters failing patches, pushing our final result to 60.9% - surp

5/5 Takeaway: pipeline order is a hyperparameter. If you're already paying for parallel rollouts, reuse them - they're relevant context, not…

Model ReleasesDGX agent

5/5 Takeaway: pipeline order is a hyperparameter. If you're already paying for parallel rollouts, reuse them - they're relevant context, not just candidate answers. Full write-up: [https://www.ai21.co

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

Model ReleasesDGX agent

arXiv:2606.04291v1 Announce Type: new Abstract: 3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains f

A couple of weeks ago we started rebuilding the 𝚑𝚏 CLI with AI agents in mind. it now detects when an agent is using it and gives clean, t…

Model ReleasesDGX agent

A couple of weeks ago we started rebuilding the 𝚑𝚏 CLI with AI agents in mind. it now detects when an agent is using it and gives clean, token-efficient output, next-command hints, and more, all desig

A New Angle on Bones: Robust Pose Estimation in X-Ray and Ultrasound

Model ReleasesDGX agent

arXiv:2606.04700v1 Announce Type: new Abstract: Measuring the angle between bone structures is a routine task in medical image analysis and provides a key quantitative parameter for diagnosis and trea

A Study of the Scale Invariant Signal to Distortion Ratio in Speech Separation with Noisy References

Model ReleasesDGX agent

arXiv:2508.14623v2 Announce Type: replace-cross Abstract: This paper examines the implications of using the Scale-Invariant Signal-to-Distortion Ratio (SI-SDR) as both evaluation and training objectiv

A Systematic Evaluation of Positional Bias in Multi-Video Summarization with MLLMs

Model ReleasesDGX agent

arXiv:2606.04596v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly used for video understanding, yet their reliability under multi-video inputs remains poorly un

Activation-Based Active Learning for In-Context Learning: Challenges and Insights

Model ReleasesDGX agent

arXiv:2606.05134v1 Announce Type: new Abstract: Deep active learning has previously been explored for LLM in-context sample selection, but not with methods that utilise recent advances in understandin

AdaKoop: Efficient Modeling of Nonlinear Dynamics from Nonstationary Data Streams with Koopman Operator Regression

Model ReleasesDGX agent

arXiv:2606.04930v1 Announce Type: cross Abstract: Real-time data analysis requires the ability to accurately and adaptively address nonlinear dynamics in a nonstationary data stream while preserving c

Adaptive Minds: Empowering Agents with LoRA-as-Tools

Model ReleasesDGX agent

arXiv:2510.15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke. We hyp

Affordance2Action: Task-Conditioned Scene-level Affordance Grounding for Real-Time Manipulation

Model ReleasesDGX agent

arXiv:2606.04172v1 Announce Type: new Abstract: Task-conditioned manipulation requires grounding instructions to task-relevant functional parts rather than object categories. This setting is scene-dep

Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents

Model ReleasesDGX agent

arXiv:2606.04874v1 Announce Type: new Abstract: Planning is central to LLM agents: before acting, an agent must decompose goals, select tools, reason over constraints, and decide when a task is infeas

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety

Model ReleasesDGX agent

arXiv:2606.04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.AI rapidly grow, concerns about unsafe human-AI interactions have intensified. This study introd

AIP: A Graph Representation for Learning and Governing Agent Skills

Model ReleasesDGX agent

arXiv:2606.04781v1 Announce Type: new Abstract: Agent Skills today consist largely of free-form prose requiring the agent to read, interpret, and re-derive how to act in every session. This imposes tw

AlgoVeri: An Aligned Benchmark for Verified Code Generation on Classical Algorithms

Model ReleasesDGX agent

arXiv:2602.09464v2 Announce Type: replace-cross Abstract: Vericoding refers to the generation of formally verified code from rigorous specifications. Recent AI models show promise in vericoding, but a

Aligning Deep Implicit Preferences by Learning to Reason Defensively

Model ReleasesDGX agent

arXiv:2510.11194v3 Announce Type: replace Abstract: Personalized alignment is crucial for enabling Large Language Models (LLMs) to engage effectively in user-centric interactions. However, current met

ALINC: Active Learning for Inductive Node Classification via Graph Sampling

Model ReleasesDGX agent

arXiv:2606.04647v1 Announce Type: new Abstract: Active learning (AL) for node classification typically focuses on selecting the most informative nodes for annotation within one or a few large graphs (

An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers

Model ReleasesDGX agent

arXiv:2606.04752v1 Announce Type: cross Abstract: Transformers consuming multi-channel scalar signals must embed C simultaneous values into one d_{ext{model}}-dimensional vector per time step. We empi

An Empirical Study of Data Scale, Model Complexity, and Input Modalities in Visual Generalization

Model ReleasesDGX agent

arXiv:2606.04409v1 Announce Type: cross Abstract: Modern deep neural networks usually have large parameter scales and nonlinear hierarchical structures, and they have achieved strong performance in co

An Open-Source Two-Stage Computer Vision Pipeline for Fine-Grained Vehicle Classification using Vision Transformers

Model ReleasesDGX agent

arXiv:2606.05149v1 Announce Type: new Abstract: Vehicle body type is a significant determinant of cyclist injury severity in overtaking crashes, yet automated tools for classifying vehicles into injur

Analysis-Driven Procedural Generation of an Engine Sound Dataset with Embedded Control Annotations

Model ReleasesDGX agent

arXiv:2603.07584v2 Announce Type: replace-cross Abstract: Computational engine sound modeling is central to the automotive audio industry, particularly for active sound design applications and virtual

And another open-weight release. Nemotron 3 Ultra has an ultra impressive capability:efficiency ratio! Design-wise, it carries forward the M…

Model ReleasesDGX agent

And another open-weight release. Nemotron 3 Ultra has an ultra impressive capability:efficiency ratio! Design-wise, it carries forward the Mamba-2-attention hybrid stack and LatentMoE introduced in th

Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna https://latent.space/p/andon @andonlabs …

Model ReleasesDGX agent

Andon Labs' Real-World AI Evals: Claude calls the FBI, AI CEOs, price cartels, Butter-Bench, & Luna https://latent.space/p/andon @andonlabs cofounders @lukaspet and @axelbacklund explain why dollar-de

← Previous
1…167168169170171…377
Next →