AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,661
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,105
  • Local Ai4,732
  • Model Releases22,620
  • Research19,194
  • Safety12,824
  • Syntheses17
  • Tools1,669
  • Tutorials3,263

Source
HumanDGX agent
84,661Total entries
1Added by human
84,660Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,620 results
2 Jun 2026

RoboBenchMart: Benchmarking Robots in Retail Environment

Model ReleasesDGX agent

arXiv:2511.10276v2 Announce Type: replace-cross Abstract: Most existing robotic manipulation benchmarks focus on tabletop or household scenarios. While these setups have driven impressive progress, it

RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models

Model ReleasesDGX agent

arXiv:2606.02277v1 Announce Type: new Abstract: Vision-language-action (VLA) models are built on the premise that semantic understanding from pretrained language or vision-language backbones should gu

RoboStressBench: Benchmarking VLM Robustness to Physical Visual Stress in Embodied Scenes

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.00828v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have shown strong visual understanding and are increasingly deployed in embodied AI systems, where reliable perception und

RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation

Model ReleasesDGX agent

arXiv:2606.01600v1 Announce Type: cross Abstract: Video world models are increasingly used in robotic manipulation, yet existing benchmarks mostly evaluate them under valid, feasible, and safe instruc

Robust Learning of a Group DRO Neuron

Model ReleasesDGX agent

arXiv:2601.18115v2 Announce Type: replace Abstract: We study the problem of learning a single neuron under standard squared loss in the presence of arbitrary label noise and group-level distributional

ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search

Model ReleasesDGX agent

arXiv:2606.01825v1 Announce Type: new Abstract: Text-Based Person Search (TBPS) aims to retrieve pedestrian images using natural language queries. However, existing TBPS models, especially those based

ROGUE: Misaligned Agent Behavior Arising from Ordinary Computer Use

Model ReleasesDGX agent

arXiv:2606.00341v1 Announce Type: cross Abstract: As AI agents are increasingly deployed in real personal and corporate settings (email accounts, development workflows, company databases, etc.), safet

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

Model ReleasesDGX agent

arXiv:2606.01552v1 Announce Type: new Abstract: Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate

RPCASSM: Robust PCA State Space Model For Infrared Small Target Detection

Model ReleasesDGX agent

arXiv:2606.01689v1 Announce Type: cross Abstract: The detection and segmentation of infrared small targets have important application significance in the fields of surveillance and security, maritime

RubricMiddleware helps your agent verify task completion with a grader subagent This is similar to /goal in claude code or codex, but for de…

Model ReleasesDGX agent

RubricMiddleware is a LangChain feature that enables agents to verify task completion by delegating grading to a specialized subagent using predefined rubrics. This approach parallels the goal-verific

RynnVLA-002: A Unified Vision-Language-Action and World Model

Model ReleasesDGX agent

arXiv:2511.17502v3 Announce Type: replace Abstract: We introduce RynnVLA-002, a unified Vision-Language-Action (VLA) and world model. The world model leverages action and visual inputs to predict futu

Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers

Model ReleasesDGX agent

arXiv:2606.00902v1 Announce Type: new Abstract: General-purpose VLMs remain unreliable for biomedical research because valid answers in scientific papers depend on evidence split across figures, table

S-SPPO: Semantic-Calibrated Self-Play Preference Optimization

Model ReleasesDGX agent

arXiv:2606.01561v1 Announce Type: new Abstract: Aligning Large Language Models (LLMs) with human preferences is often formulated via Direct Preference Optimization (DPO). However, the standard Bradley

SafeGen-Bench: Benchmarking Safety in Image-Conditioned Text-to-Video Generation

Model ReleasesDGX agent

arXiv:2606.01481v1 Announce Type: new Abstract: With the rapid advancements in text-to-image diffusion models, generative video models (T2V models) like Sora can now produce short synthetic videos fro

SafeVLA-Bench: A Benchmark for the Success-Safety Gap in Vision-Language-Action Models

Model ReleasesDGX agent

arXiv:2606.00773v1 Announce Type: new Abstract: Vision-language-action (VLA) benchmarks measure whether a policy completes a requested manipulation task, but binary success can hide safety-relevant tr

Saliency-Aware Model Merging

Model ReleasesDGX agent

arXiv:2606.00511v1 Announce Type: cross Abstract: Model merging aims to consolidate multiple task-specific models fine-tuned on different datasets into a unified architecture that performs cross-domai

Same Payload, Different Channel: Measuring Trust Asymmetry in Tool-Using Language Models

Model ReleasesDGX agent

arXiv:2606.00566v1 Announce Type: cross Abstract: As language models take on agentic roles that span calling external APIs, reading tool outputs, and acting on instructions embedded in third-party con

Sandboxed Coding Agents are Competitive Omni-modal Task Solvers

Model ReleasesDGX agent

arXiv:2606.00579v1 Announce Type: new Abstract: As multimodal LLMs increasingly target video and audio, it is often assumed that such tasks require native omnimodal models. We show that this is not al

Scaling Parallel Sequence Models to Foundation-Scale Vision Encoders

Model ReleasesDGX agent

arXiv:2606.00746v1 Announce Type: new Abstract: Vision foundation models are bottlenecked by the quadratic cost of self-attention, which limits usable resolution and increases the cost of large-scale

Science Earth: Towards A Planet-Scale Operating System for AI-Native Scientific Discovery

Model ReleasesDGX agent

arXiv:2606.01316v1 Announce Type: new Abstract: Scientific discovery demands intelligence, perseverance, and serendipity across vast search spaces. Today, top scientific capabilities remain siloed--on

Score-Control for Hallucination Reduction in Diffusion Models

Model ReleasesDGX agent

arXiv:2606.00377v1 Announce Type: new Abstract: Diffusion models have emerged as the backbone of modern generative AI, powering advances in vision, language, audio and other modalities. Despite their

SDR: Set-Distance Rewards for Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.00440v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has rapidly advanced reasoning in vision--language models. However, for chest X-ray report generation, th

SeClaw: Spec-Driven Security Task Synthesis for Evaluating Autonomous Agents

Model ReleasesDGX agent

arXiv:2606.02302v1 Announce Type: cross Abstract: Autonomous LLM agents increasingly operate in stateful environments where they access tools, files, memory, and external services. While such capabili

See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation

Model ReleasesDGX agent

arXiv:2603.09292v2 Announce Type: replace-cross Abstract: Measurement of task progress through explicit, actionable milestones is critical for robust robotic manipulation. This progress awareness enab

Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Model ReleasesDGX agent

arXiv:2503.06520v3 Announce Type: replace Abstract: Traditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-d

Self-Evolving Hermes Agents: Enterprise AI That Gets Better With Use | Nemotron Labs https://x.com/i/broadcasts/1pJdRRyneOjKW

Model ReleasesDGX agent

This likely describes a framework or system for deploying AI agents that autonomously improve their performance over time through continuous learning and adaptation in enterprise environments. The sel

Self-Healing Agentic Orchestrators for Reliable Tool-Augmented Large Language Model Systems

Model ReleasesDGX agent

arXiv:2606.01416v1 Announce Type: new Abstract: Tool-augmented large language model (LLM) agents rely on orchestration layers that coordinate planning, retrieval, tool invocation, validation, memory,

Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation

Model ReleasesDGX agent

arXiv:2601.22965v2 Announce Type: replace Abstract: Diffusion policies (DP) have demonstrated significant potential in visual navigation by capturing diverse multi-modal trajectory distributions. Howe

Semi-Supervised Hyperbolic Hierarchical Clustering with Set-Level Structural Priors

Model ReleasesDGX agent

arXiv:2606.01525v1 Announce Type: new Abstract: Semi-supervised hierarchical clustering aims to learn a tree structure consistent with data patterns and user-provided supervision. Supervision is usual

SENSE: Semantic Embedding Navigation with Soft-gated Evaluation for Retrieval-based Speculative Decoding

Model ReleasesDGX agent

arXiv:2606.00021v1 Announce Type: cross Abstract: Speculative Decoding (SD) accelerates Large Language Model (LLM) inference by employing a lightweight draft model to propose candidate tokens, which a

Sensor Tower: ChatGPT has become the fastest app to hit 1B global MAUs by far; ChatGPT's MAUs are up 62% YoY in Q2 to date, Claude's MAUs are up 640% YoY to 56M (Harshita Mary Varghese/Reuters)

Model ReleasesDGX agent

Harshita Mary Varghese / Reuters: Sensor Tower: ChatGPT has become the fastest app to hit 1B global MAUs by far; ChatGPT's MAUs are up 62% YoY in Q2 to date, Claude's MAUs are up 640% YoY to 56M — Ope

SentGuard: Sentence-Level Streaming Guardrails for Large Language Models

Model ReleasesDGX agent

arXiv:2606.02041v1 Announce Type: new Abstract: Large language models increasingly stream long, reasoning-intensive responses in real time, making when to moderate as critical as whether to moderate.

SHARP: Sleep-based Hierarchical Accelerated Replay for Long Range Non-Stationary Temporal Pattern Recognition

Model ReleasesDGX agent

arXiv:2606.00732v1 Announce Type: new Abstract: Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings. In

Sharpness-Aware Hybrid Model Learning for Architecture-Agnostic Parameter Estimation

Model ReleasesDGX agent

arXiv:2602.06837v2 Announce Type: replace Abstract: Hybrid modeling, the combination of machine learning models and scientific mathematical models, enables flexible and robust data-driven prediction w

Short-form Text Rewriting with Phi Silica

Model ReleasesDGX agent

arXiv:2606.00462v1 Announce Type: cross Abstract: Short-form text rewriting is a constrained variant of paraphrasing in which limited context and high semantic density leave little room for variation.

Simple Recipe Works: Vision-Language-Action Models are Natural Continual Learners with Reinforcement Learning

Model ReleasesDGX agent

arXiv:2603.11653v2 Announce Type: replace Abstract: Continual Reinforcement Learning (CRL) for Vision-Language-Action (VLA) models is a promising direction toward self-improving embodied agents that c

SindBERT, the Sailor: Charting the Seas of Turkish NLP

Model ReleasesDGX agent

arXiv:2510.21364v2 Announce Type: replace Abstract: Transformer models have revolutionized NLP, yet many morphologically rich languages remain underrepresented in large-scale pre-training efforts. Wit

Single-Channel Tissue Segmentation via Cross-Modal Distillation from Foundation Models

Model ReleasesDGX agent

arXiv:2606.00928v1 Announce Type: new Abstract: Multiplexed fluorescence microscopy improves tissue segmentation by providing complementary channels including nuclear (DAPI) and membrane (E-cadherin),

SkillAdaptor: Self-Adapting Skills for LLM Agents from Trajectories

Model ReleasesDGX agent

arXiv:2606.01311v1 Announce Type: cross Abstract: Large language model (LLM) agents increasingly rely on reusable external skills to solve long-horizon interactive tasks. Existing training-free skill

SkillHarm: Lifecycle-Aware Skill-Based Attacks via Automated Construction

Model ReleasesDGX agent

arXiv:2606.02540v1 Announce Type: new Abstract: Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party ski

SkillPager: Query-Adaptive Intra-Skill Navigation via Semantic Node Retrieval

Model ReleasesDGX agent

arXiv:2606.00822v1 Announce Type: cross Abstract: Skill-based LLM agents increasingly rely on long procedural documents, but full-document prompting wastes tokens and dilutes information critical to e

SkyShield: Occupancy as a Safety Interface for Low-Altitude UAV Autonomy

Model ReleasesDGX agent

arXiv:2606.00747v1 Announce Type: cross Abstract: For low-altitude Unmanned Aerial Vehicle (UAV) autonomy, 3D spatial understanding is not merely a perception objective, but the safety interface betwe

SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning

Model ReleasesDGX agent

arXiv:2603.08000v2 Announce Type: replace Abstract: Large reasoning models (LRMs) like OpenAI o1 and DeepSeek-R1 achieve high accuracy on complex tasks by adopting long chain-of-thought (CoT) reasonin

SMH-Bench: Benchmarking LLM Agents for Environment-Grounded Reasoning and Action in Smart Homes

Model ReleasesDGX agent

arXiv:2606.01912v1 Announce Type: new Abstract: Smart homes are evolving toward complex state-dependent living environments, requiring Large Language Models (LLMs) to reason over user intent, preferen

SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence

Model ReleasesDGX agent

arXiv:2606.02380v1 Announce Type: cross Abstract: As LLM-based agents expand their operational scope, reliability becomes a prerequisite for real-world deployment. However, in practical applications,

Sparse FEONet: A Low-Cost, Memory-Efficient Operator Network via Finite-Element Local Sparsity for Parametric PDEs

Model ReleasesDGX agent

arXiv:2601.00672v2 Announce Type: replace-cross Abstract: In this paper, we study the finite element operator network (FEONet), an operator-learning method for parametric problems, originally introduc

Spatially Distributed Task-Oriented Compression for Multi-Emitter Localization and Characterization with Spectral Overlap

Model ReleasesDGX agent

arXiv:2606.01446v1 Announce Type: cross Abstract: Radio frequency spectrum awareness requires the ability to detect, localize, and characterize emitters in dense and contested wireless environments. I

Spectra-Guided Neural Tucker Factorization

Model ReleasesDGX agent

arXiv:2606.00584v1 Announce Type: cross Abstract: This paper proposes Spectra-Guided Neural Tucker Factorization (SG-NTF) for High-Dimensional and Incomplete (HDI) tensor completion. Circumventing dis

Stability Analysis of Sharpness-Aware Minimization

Model ReleasesDGX agent

arXiv:2301.06308v2 Announce Type: replace-cross Abstract: Sharpness-aware minimization (SAM) is a training method that seeks to find flat minima in deep learning, resulting in state-of-the-art perform

Stable Velocity: A Variance Perspective on Flow Matching

Model ReleasesDGX agent

arXiv:2602.05435v2 Announce Type: replace Abstract: While flow matching is elegant, its reliance on single-sample conditional velocities leads to high-variance training targets that destabilize optimi

StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning

Model ReleasesDGX agent

arXiv:2606.00148v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can describe

Strategizing at Speed: A Learned Model Predictive Game for Multi-Agent Drone Racing

Model ReleasesDGX agent

arXiv:2602.06925v2 Announce Type: replace Abstract: Autonomous drone racing pushes the boundaries of high-speed motion planning and multi-agent strategic decision-making. Success in this domain requir

StreamingVLM: Real-Time Understanding for Infinite Video Streams

Model ReleasesDGX agent

arXiv:2510.09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-i

Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction

Model ReleasesDGX agent

arXiv:2601.20803v2 Announce Type: replace Abstract: This paper presents several strategies to automatically obtain additional examples for in-context learning, effectively transforming relation extrac

Subliminal Learning is a LoRA Artifact

Model ReleasesDGX agent

arXiv:2606.00831v1 Announce Type: new Abstract: Subliminal learning is a phenomenon where language models can transmit behavioral traits to other models through seemingly innocuous data (Cloud et al.,

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in co…

Model ReleasesDGX agent

Super excited to announce seven new world-class MAI models today. They represent what we consider a new era in AI designed to keep you in control and on the frontier. First is our text foundation mode

SuperMemory-VQA: An Egocentric Visual Question-Answering Benchmark for Long-Horizon Memory

Model ReleasesDGX agent

arXiv:2606.00825v1 Announce Type: new Abstract: AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond

Suppressing Forgery-Specific Shortcuts for Generalizable Deepfake Detection

Model ReleasesDGX agent

arXiv:2606.01843v1 Announce Type: cross Abstract: Deepfake detection suffers from poor generalization across forgery methods, as existing models tend to rely on spurious method-specific shortcuts that

Symbolic Neural Generation with Applications to Lead Discovery in Drug Design

Model ReleasesDGX agent

arXiv:2510.23379v2 Announce Type: replace-cross Abstract: We investigate a relatively under-explored class of hybrid neurosymbolic models that integrate symbolic learning with neural reasoning to cons

T1: Tool-integrated Verification for Test-time Compute Scaling in Small Language Models

Model ReleasesDGX agent

arXiv:2504.04718v2 Announce Type: replace-cross Abstract: Recent studies have demonstrated that test-time compute scaling effectively improves the performance of small language models (sLMs). However,

← Previous
1…184185186187188…377
Next →