AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,612 results
29 May 2026

Probabilistic bias adjustment of seasonal forecasts using generative machine learning: A case study of Arctic sea ice predictions

Model ReleasesDGX agent

arXiv:2605.29172v1 Announce Type: new Abstract: Seasonal climate predictions support planning and risk management by offering early information of the most likely-to-occur climate conditions in the co

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

Model ReleasesDGX agent

arXiv:2605.30284v1 Announce Type: new Abstract: Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks ha

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases
DGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data

Model ReleasesDGX agent

arXiv:2508.15180v3 Announce Type: replace Abstract: High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

Model ReleasesDGX agent

arXiv:2605.30280v1 Announce Type: cross Abstract: Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented cap

RAISE: RAG Design as an Architecture Search Problem

Model ReleasesDGX agent

arXiv:2605.30029v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

Model ReleasesDGX agent

arXiv:2605.29579v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinat

Realistic honeypot evaluations for scheming propensity

Model ReleasesDGX agent

arXiv:2605.29729v1 Announce Type: new Abstract: We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.29114v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

Model ReleasesDGX agent

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

Model ReleasesDGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

Model ReleasesDGX agent

arXiv:2603.05488v4 Announce Type: replace-cross Abstract: We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer,

ReasonOps: Operator Segmentation for LLM Reasoning Traces

Model ReleasesDGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

Model ReleasesDGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

Model ReleasesDGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

Relational Rank Geometry in Transformers: Detecting and Steering Hidden-State Relation Frames

Model ReleasesDGX agent

arXiv:2605.29634v1 Announce Type: new Abstract: Transformer hidden states are often interpreted through local or low-order objects: neurons, sparse features, attention heads, residual-stream direction

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

Model ReleasesDGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

REPOT: Recoverable Program-of-Thought via Checkpoint Repair

Model ReleasesDGX agent

arXiv:2605.30052v1 Announce Type: cross Abstract: One-shot Program-of-Thought (PoT) emits a Python program that prints a primitive-action plan; a single invalid action silently invalidates the traject

Representation Unlearning: Forgetting through Information Compression

Model ReleasesDGX agent

arXiv:2601.21564v2 Announce Type: replace Abstract: Machine unlearning seeks to remove the influence of specific training data from a model, a need driven by privacy regulations and robustness concern

Rethinking Literature Search Evaluation: Deep Research Helps, and Human Citation Lists Are Not a Ground Truth

Model ReleasesDGX agent

arXiv:2605.29234v1 Announce Type: new Abstract: We study large-scale literature search from two complementary angles: improving the retrieval pipeline, and stress-testing the human reference list as a

Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting

Model ReleasesDGX agent

arXiv:2605.29401v1 Announce Type: new Abstract: Time-Series Foundation Models (TSFMs) excel at zero-shot unimodal forecasting using numerical data, but unlike LLMs they cannot consume multimodal, non-

Return-to-Go Is More Than a Number: Q-Guided Alignment for Return-Conditioned Supervised Learning

Model ReleasesDGX agent

arXiv:2605.29028v1 Announce Type: cross Abstract: Conditioned Sequence Models (CSMs) learn policies by treating return-to-go (RTG) as a control signal. However, existing CSMs often treat the RTGs as s

RHO: Robust Holistic OSM-Based Metric Cross-View Geo-Localization

Model ReleasesDGX agent

arXiv:2603.27758v2 Announce Type: replace Abstract: Metric Cross-View Geo-Localization (MCVGL) aims to estimate the 3-DoF camera pose (position and heading) by matching ground and satellite images. In

RightNow-Arabic-0.5B-Turbo: An Open Sub-1B Arabic Language Model via Vocabulary Injection and Edge-First Deployment

Model ReleasesDGX agent

arXiv:2605.28827v1 Announce Type: new Abstract: Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arabic as an afterthought (Qwen2.5-0.5B, Falcon-H1-0.5B)

Risk-averse Fair Multi-class Classification

Model ReleasesDGX agent

arXiv:2509.05771v2 Announce Type: replace-cross Abstract: We develop a new classification framework based on the theory of coherent risk measures and systemic risk. The proposed approach is suitable f

RoboWits: Unexpected Challenges for Robotic Creative Problem Solving

Model ReleasesDGX agent

arXiv:2605.30326v1 Announce Type: cross Abstract: The ability to reason, adapt, and creatively solve problems under unexpected challenges is essential for robots operating in real-world environments.

Robust and Efficient Guardrails with Latent Reasoning

Model ReleasesDGX agent

arXiv:2605.29068v1 Announce Type: new Abstract: Maintaining the safety of large language models (LLMs) is crucial as they are increasingly deployed in real-world applications. Existing safety guardrai

S-MARC: Causal Streaming Reasoning for Full-Duplex Conversational Behavior Modeling

Model ReleasesDGX agent

arXiv:2602.11065v2 Announce Type: replace-cross Abstract: Human conversation is organized by an implicit chain of thought and manifests as temporally structured conversational behaviors. Capturing thi

SAAS: Self-Aware Reinforcement Learning for Over-Search Mitigation in Agentic Search

Model ReleasesDGX agent

arXiv:2605.29796v1 Announce Type: new Abstract: Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these syste

SafeSearch: Automated Red-Teaming of LLM-Based Search Agents

Model ReleasesDGX agent

arXiv:2509.23694v5 Announce Type: replace Abstract: Search agents connect LLMs to the Internet, enabling them to access broader and more up-to-date information. However, this also introduces a new thr

SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation

Model ReleasesDGX agent

arXiv:2507.09266v2 Announce Type: replace Abstract: Gloss-free Sign Language Translation (SLT) has advanced rapidly, achieving strong performances without relying on gloss annotations. However, these

Salesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days…

Model ReleasesDGX agent

Salesforce published a detailed writeup on going agentic with Claude Code. A couple things jumped out. A migration they'd scoped at 231 days shipped in 13. One PR delivered 21 endpoints at 100% test c

Same Question, Different Source, Different Answer: Auditing Source-Dependence in Medical Multi-Source RAG

Model ReleasesDGX agent

arXiv:2605.29084v1 Announce Type: cross Abstract: A retrieval-augmented generation (RAG) system deployed over a multi-author institutional corpus can give a different answer to the same question depen

Sample-Efficient Diffusion-based Reinforcement Learning with Critic Guidance

Model ReleasesDGX agent

arXiv:2605.30056v1 Announce Type: cross Abstract: Recent advances in reinforcement learning (RL) have achieved great successes by leveraging the multimodality and exploration capability of diffusion p

Scaling Laws for Agent Harnesses via Effective Feedback Compute

Model ReleasesDGX agent

arXiv:2605.29682v1 Announce Type: new Abstract: Agent harnesses increasingly determine the performance of language-model systems by deciding how models call tools, receive feedback, verify intermediat

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

Model ReleasesDGX agent

arXiv:2605.29358v1 Announce Type: new Abstract: We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

Model ReleasesDGX agent

arXiv:2605.29059v1 Announce Type: cross Abstract: Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studi

SciIntBench: Measuring LLM Compliance with Research Integrity Norms Under Adversarial Framing

Model ReleasesDGX agent

arXiv:2605.29468v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to support scientific work, but it is unclear whether they uphold responsible conduct of research (

SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

Model ReleasesDGX agent

arXiv:2512.15374v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly deployed in environments that generate massive, dynamic contexts. However, a critical bottleneck

SDF-Net: Structure-Aware Disentangled Feature Learning for Opticall-SAR Ship Re-identification

Model ReleasesDGX agent

arXiv:2603.12588v2 Announce Type: replace Abstract: Cross-modal ship re-identification (ReID) between optical and synthetic aperture radar (SAR) imagery is fundamentally challenged by the severe radio

Selection Hyper-heuristics Can Automatically Adjust the Learning Period to Optimally Solve Pseudo-Boolean Problems

Model ReleasesDGX agent

arXiv:2605.29916v1 Announce Type: cross Abstract: The Random Gradient hyper-heuristic was recently shown to be able to learn the optimal neighbourhood size when optimizing the LeadingOnes benchmark vi

Selective QA over Conflicting Multi-Source Personal Memory: A Diagnostic Testbed and Method Comparison

Model ReleasesDGX agent

arXiv:2605.30087v1 Announce Type: new Abstract: Emerging personal AI agents are moving toward persistent, multi-source memory. This creates an evaluation problem: systems must decide how to use confli

Semantic and Visual Evidence for Efficient Long-Video Reasoning: A Solution for the HD-EPIC VQA Challenge

Model ReleasesDGX agent

arXiv:2605.29402v1 Announce Type: cross Abstract: Understanding long-form egocentric videos remains challenging for multimodal large language models (MLLMs) due to limited context length and insuffici

Sequential Physics-Constrained Neural Operator Forward Modeling for the extit{Norne} Reservoir System

Model ReleasesDGX agent

arXiv:2605.28909v1 Announce Type: new Abstract: We develop a comprehensive mathematical and computational framework for sequential surrogate modeling of three-phase black-oil reservoir dynamics using

SERC: LDPC-Inspired Semantic Error Correction for Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2605.28837v1 Announce Type: cross Abstract: While Large Language Models (LLMs) have demonstrated remarkable capabilities, their reliability is significantly compromised by hallucinations. Existi

Skill-Pro: Learning Reusable Skills from Experience via Non-Parametric PPO for LLM Agents

Model ReleasesDGX agent

arXiv:2602.01869v3 Announce Type: replace Abstract: LLM-driven agents excel at sequential decision-making but often rely on on-the-fly reasoning, re-deriving solutions even in recurring scenarios. Thi

SkillsInjector: Dynamic Skill Context Construction for LLM Agents

Model ReleasesDGX agent

arXiv:2605.29794v1 Announce Type: new Abstract: LLM agents now draw on growing skill libraries to handle complex tasks. However, injecting more skills does not always improve task completion and can e

SLAD : Shared LoRA Adapters for Task Specific Distillation

Model ReleasesDGX agent

arXiv:2605.29726v1 Announce Type: new Abstract: In the context of resource-constrained environments such as embedded systems, adapting reduced-size foundation models to downstream tasks has become inc

Small Agent Group is the Future of Digital Health

Model ReleasesDGX agent

arXiv:2602.08013v2 Announce Type: replace Abstract: The rapid adoption of large language models (LLMs) in digital health has been driven by a 'scaling-first' philosophy, i.e., the assumption that clin

SMolLM: Small Language Models Learn Small Molecular Grammar

Model ReleasesDGX agent

arXiv:2605.06322v2 Announce Type: replace Abstract: Language models for molecular design have scaled to hundreds of millions of parameters, yet how they learn chemical grammar is poorly understood. We

Some fun Gemini Omni use cases from the community 🧵👇

Model ReleasesDGX agent

This X thread from Google AI showcases community-created use cases and applications of Gemini Omni, Google's multimodal AI model. The post likely highlights practical and creative examples of how user

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones?

Model ReleasesDGX agent

arXiv:2605.30329v1 Announce Type: new Abstract: Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. How

Stable-Layers: Fine-Tuning Image Layer Decomposition Models with VLM-Scored Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.30257v1 Announce Type: new Abstract: We present Stable-Layers, a reinforcement learning framework that eliminates the need for paired supervision by fine-tuning a pretrained layer decomposi

STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments

Model ReleasesDGX agent

arXiv:2605.29324v1 Announce Type: new Abstract: Mobile GUI agents excel at immediate reactive control but frequently fail in realistic, long-horizon tasks that require memory. This failure stems from

Strengthening societal resilience with Rosalind Biodefense

Model ReleasesDGX agent

OpenAI launches Rosalind Biodefense, expanding trusted access to GPT-Rosalind for vetted developers and U.S. government partners advancing biodefense, public health, and pandemic preparedness through

Striding Across Reynolds Numbers: Representation Geometry in Neural PDE Generalisation

Model ReleasesDGX agent

arXiv:2605.30112v1 Announce Type: new Abstract: Cross-Reynolds generalisation in neural PDE solvers remains poorly characterised. On the canonical forced 2D Navier-Stokes benchmark, a trained Fourier

Structure-Aware Text Recognition for Ancient Greek Critical Editions

Model ReleasesDGX agent

arXiv:2603.02803v2 Announce Type: replace Abstract: Recent advances in visual language models (VLMs) have transformed end-to-end document understanding. However, their ability to interpret the complex

SURGENT: A Surgical Multi-Agent Assistance System Across the Perioperative Workflow

Model ReleasesDGX agent

arXiv:2605.29368v1 Announce Type: cross Abstract: The intricate nature of modern surgical care necessitates intelligent systems that can synthesize extensive patient records, support collaborative dec

SwInception -- Local Attention Meets Convolutions

Model ReleasesDGX agent

arXiv:2605.29954v1 Announce Type: new Abstract: Sparse vision transformers have gained popularity as efficient encoders for medical volumetric segmentation, with Swin emerging as a prominent choice. S

TAE: Target-aware enhancer for nighttime UAV tracking

Model ReleasesDGX agent

arXiv:2605.29558v1 Announce Type: new Abstract: Severe image degradation under low-light nighttime conditions constitutes a core bottleneck preventing all-day applications for UAV-based single object

← Previous
1…195196197198199…377
Next →