AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,648
  • Agents7,273
  • Applications5,201
  • Concepts5
  • Hardware1,758
  • Industry6,104
  • Local Ai4,732
  • Model Releases22,612
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,669
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,648Total entries
1Added by human
84,647Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

NICE: A Theory-Grounded Diagnostic Benchmark for Social Intelligence of LLMs

DGX agent

arXiv:2605.29685v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly applied in social contexts such as emotional companionship and customer service, measuring their social

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

DGX agent

arXiv:2605.30120v1 Announce Type: cross Abstract: Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-le

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems

DGX agent

arXiv:2605.29676v1 Announce Type: new Abstract: Large language models in Agentic AI systems consume tool schemas and execution results and emit tool invocations as structured data. The default languag

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OISD: On-Policy Internal Self-Distillation of Language Models

DGX agent

arXiv:2605.29089v1 Announce Type: cross Abstract: Recent reinforcement learning (RL) post-training approaches primarily optimize the final output policy using sparse outcome-level rewards, while large

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OmniMatBench: A Human-Calibrated Multimodal Reasoning Benchmark Across 19 Materials Science Subfields

DGX agent

arXiv:2605.29833v1 Announce Type: new Abstract: As multimodal language models play an increasingly important role in scientific research, materials science offers a critical testbed due to its interdi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources

DGX agent

arXiv:2605.29250v1 Announce Type: cross Abstract: Real-world information needs require access to structurally diverse knowledge sources, from unstructured text and relational tables to knowledge graph

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

On-Policy Replay for Continual Supervised Fine-Tuning

DGX agent

arXiv:2605.29495v1 Announce Type: new Abstract: Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

On the Construction and Implications of Low-Loss Valleys in LoRA-based Bayesian Inference

DGX agent

arXiv:2605.29580v1 Announce Type: new Abstract: While parameter-efficient fine-tuning methods like low-rank adaptation (LoRA) are standard for large language models, principled estimation of epistemic

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

OOD-GraphLLM: Graph Large Language Model for Out-of-Distribution Generalized Drug Synergy Prediction

DGX agent

arXiv:2605.30247v1 Announce Type: new Abstract: Drug synergy prediction (DSP) aims to identify efficacious drug combinations under various cellular contexts with different targets. However, the contin

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

OpenClawBench: Benchmarking Process-side Anomalies in Real-world Agent Execution Trajectories

DGX agent

arXiv:2605.29253v1 Announce Type: new Abstract: Task success can hide process anomalies in real-world agent executions. An agent may pass the final task oracle while still accumulating unresolved ambi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Opir: Efficient Multi-Task Safety Classification for Toxicity, Jailbreaks, Hate Speech, and Harmful Content

DGX agent

arXiv:2605.29659v1 Announce Type: cross Abstract: Real-time safety filtering for large language model (LLM) applications requires classifiers that can detect unsafe prompts, toxic language, jailbreak

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Optimizing Latent Representations for Robust Building Damage Assessment Onboard Earth Observation Satellites

DGX agent

arXiv:2605.29575v1 Announce Type: new Abstract: Rapid identification of damaged buildings after natural disasters or on war areas is crucial to support emergency response and prioritize interventions.

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

OptSkills: Learning Generalizable Optimization Skills from Problem Archetypes via Cluster-Based Distillation

DGX agent

arXiv:2605.29829v1 Announce Type: new Abstract: Leveraging Large Language Models (LLMs) to automatically formulate and solve optimization problems from natural language has emerged as an efficient par

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Orthogonal Concept Erasure for Diffusion Models

DGX agent

arXiv:2605.28902v1 Announce Type: new Abstract: Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face signifi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

OVA-IB: One vs All Information Bottleneck for Multi-Modal Alignment

DGX agent

arXiv:2605.29900v1 Announce Type: new Abstract: Contrastive learning is effective for aligning paired views or modalities, but alignment beyond two modalities remains non-trivial and comparatively und

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Overcoming Forgetting in LLM Fine-Tuning with Evolution Strategies

DGX agent

arXiv:2605.30148v1 Announce Type: cross Abstract: Evolution Strategies (ES) has recently emerged as a competitive alternative to reinforcement learning (RL) for large language model (LLM) fine-tuning,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Parallax: Parameterized Local Linear Attention for Language Modeling

DGX agent

arXiv:2605.29157v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remain

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Parameter-Efficient Subspace Decoupling ViT for Mitigating Multi-Task Negative Transfer in Histological Scoring

DGX agent

arXiv:2605.29852v1 Announce Type: new Abstract: Histological scoring is essential for diagnosing Non-Alcoholic Fatty Liver Disease (NAFLD), yet its automation remains challenging due to the high annot

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

ParaTool: Shifting Tool Representations from Context to Parameters

DGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Personalized Turn-Level User Conversation Satisfaction Benchmark

DGX agent

arXiv:2605.29711v1 Announce Type: cross Abstract: User satisfaction with AI assistants is highly personalized: the same response may satisfy one user but disappoint another depending on what each user

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PhAIL: A Real-Robot VLA Benchmark and Distributional Methodology

DGX agent

arXiv:2605.29710v1 Announce Type: new Abstract: Real-world evaluation of vision-language-action (VLA) policies still rests on binary success rate at a fixed timeout with N le 25 rollouts per condition

model-releasesarxiv-cs-ro
29 May 2026
Model Releases

PhoneWorld: Scaling Phone-Use Agent Environments

DGX agent

arXiv:2605.29486v1 Announce Type: cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale. Ex

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Physics Is All You Need? A Case Study in Physicist-Supervised AI Development of Scientific Software

DGX agent

arXiv:2605.30353v1 Announce Type: new Abstract: Are AI agents tools, co-authors, or researchers? We present a quantified case study (N=1): a physicist supervising an AI coding agent (Claude Code, Sonn

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Planning with the Views via Scene Self-Exploration

DGX agent

arXiv:2605.29563v1 Announce Type: new Abstract: Can VLMs predict how each camera move changes the view, and plan many such moves ahead? We call this capability view planning, requiring (1)understandin

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Pocket-Dentist: On-Device Dental Image Understanding via Efficient Multimodal Large Language Models

DGX agent

arXiv:2605.29299v1 Announce Type: cross Abstract: Evaluations of dental vision-language models remain fragmented across datasets, task definitions and metrics, and often ignore their computational cos

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PokerSkill: LLMs Can Play Expert-Level Poker without Training or Solvers

DGX agent

arXiv:2605.30094v1 Announce Type: new Abstract: Poker is a landmark challenge for artificial intelligence. The dominant approach relies on equilibrium solvers built on counterfactual regret minimizati

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

DGX agent

arXiv:2605.29815v1 Announce Type: new Abstract: The growing number of submitted papers has motivated the exploration of Large Language Models (LLMs) as a means to support and augment the peer review p

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Pre-Registering the Detectable Effect: A Paired-MDE Budget for 4-bit Quantization Benchmarks, with a Pilot Audit

DGX agent

arXiv:2605.28873v1 Announce Type: new Abstract: This is a planning-method note with an unpaired pilot audit. We adapt the classical paired-binary sample-size calculation (Miettinen, 1968) to quantizat

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Predicting Causal Effects from Natural Language Queries using Structured Representations

DGX agent

arXiv:2605.29631v1 Announce Type: cross Abstract: Randomized controlled trials are a cornerstone of medicine and the social sciences as they enable reliable estimates of causal effects. However, they

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Prescribe-then-Select: Adaptive Policy Selection for Contextual Stochastic Optimization

DGX agent

arXiv:2509.08194v2 Announce Type: replace Abstract: We address the problem of policy selection in contextual stochastic optimization (CSO), where covariates are available as contextual information and

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Prioritize the Process, Not Just the Outcome: Rewarding Latent Thought Trajectories Improves Reasoning in Looped Language Models

DGX agent

arXiv:2602.10520v3 Announce Type: replace Abstract: Looped Language Models (LoopLMs) perform multi-step latent reasoning prior to token generation and outperform conventional LLMs on reasoning benchma

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Probabilistic bias adjustment of seasonal forecasts using generative machine learning: A case study of Arctic sea ice predictions

DGX agent

arXiv:2605.29172v1 Announce Type: new Abstract: Seasonal climate predictions support planning and risk management by offering early information of the most likely-to-occur climate conditions in the co

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

ProjectionBench: Evaluating Scientific Hypothesis Generation in LLMs Under Progressive Information Disclosure

DGX agent

arXiv:2605.30284v1 Announce Type: new Abstract: Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks ha

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PTCG-Bench: Can LLM Agents Master Pokemon Trading Card Game?

DGX agent

arXiv:2605.29653v1 Announce Type: new Abstract: Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require sim

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

PuzzleClone: A DSL-Powered Framework for Synthesizing Verifiable Data

DGX agent

arXiv:2508.15180v3 Announce Type: replace Abstract: High-quality mathematical and logical datasets with verifiable answers are essential for strengthening the reasoning capabilities of large language

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

DGX agent

arXiv:2605.30280v1 Announce Type: cross Abstract: Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented cap

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

RAISE: RAG Design as an Architecture Search Problem

DGX agent

arXiv:2605.30029v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) systems expose numerous design choices spanning query rewriting, chunking, retrieval depth, reranking, and context

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ReactBench: A Cause-Driven Benchmark for Multimodal Hallucination via Systematic Evaluation

DGX agent

arXiv:2605.29579v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) have achieved rapid progress in vision-language understanding, they remain prone to multimodal hallucinat

model-releasesarxiv-cs-cv
29 May 2026
Model Releases

Realistic honeypot evaluations for scheming propensity

DGX agent

arXiv:2605.29729v1 Announce Type: new Abstract: We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

ReasonBreak: Probing Vulnerabilities in Reasoning-Enabled Vision-Language-Action Models for Autonomous Driving

DGX agent

arXiv:2605.29114v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models with integrated reasoning have been proposed for end-to-end autonomous driving, assuming a tight coupling between

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

DGX agent

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Reasoning-preserved Efficient Distillation of Large Language Models via Activation-aware Initialization

DGX agent

arXiv:2605.29327v1 Announce Type: new Abstract: Efficient Distillation (EDistill) compresses large language models (LLMs) by structured pruning parameters and tuning lightweight modules with high trai

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Reasoning Theater: Disentangling Model Beliefs from Chain-of-Thought

DGX agent

arXiv:2603.05488v4 Announce Type: replace-cross Abstract: We provide evidence of performative chain-of-thought (CoT) in reasoning models, where a model becomes strongly confident in its final answer,

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

ReasonOps: Operator Segmentation for LLM Reasoning Traces

DGX agent

arXiv:2605.29192v1 Announce Type: new Abstract: Chain-of-thought traces from large reasoning models can span tens of thousands of tokens, yet we lack a vocabulary for describing their internal structu

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Recovering Diversity Without Losing Alignment: A DPO Recipe for Post-Trained LLMs

DGX agent

arXiv:2605.30021v1 Announce Type: new Abstract: Many open-ended instructions have multiple valid answers that users can benefit from seeing, but post-training often narrows an LLM's output space towar

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Redundant or Necessary? A Benchmark for Detecting Redundant Steps in Agent Trajectories

DGX agent

arXiv:2605.29893v1 Announce Type: new Abstract: LLM-based agents have demonstrated strong capabilities in solving complex tasks through multi-step reasoning and tool use. However, existing evaluation

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Relational Rank Geometry in Transformers: Detecting and Steering Hidden-State Relation Frames

DGX agent

arXiv:2605.29634v1 Announce Type: new Abstract: Transformer hidden states are often interpreted through local or low-order objects: neurons, sparse features, attention heads, residual-stream direction

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Relevance as a Vulnerability: How Web Retrieval Degrades Safety Alignment in LLM Agents

DGX agent

arXiv:2605.29224v1 Announce Type: cross Abstract: AI agents augment large language models with external tools such as web retrieval, enabling grounded and up-to-date responses. However, incorporating

model-releasesarxiv-cs-ai
29 May 2026
← Previous
1…182183184185186…361
Next →