AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

DGX agent

arXiv:2606.28589v1 Announce Type: new Abstract: Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and 'Wait' prompts, primarily encourage models to think mor

model-releasesarxiv-cs-ai
30 Jun 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages

DGX agent

arXiv:2606.28715v1 Announce Type: cross Abstract: While AI development and evaluation for Southeast Asia (SEA) has grown rapidly, agent capabilities in regional languages are still poorly understood d

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Selective Memory Retention for Long-Horizon LLM Agents

DGX agent

arXiv:2606.29178v1 Announce Type: new Abstract: When does retention matter for memory-augmented LLM agents? We study this with TraceRetain, a lightweight framework for bounded external memory in froze

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Self-Supervised Theorem Discovery in a Formal Axiomatic System

DGX agent

arXiv:2606.28747v1 Announce Type: new Abstract: Recent artificial intelligence (AI) systems have shown remarkable progress in mathematical reasoning. Many existing approaches, including large language

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Semantic-Driven Scale and Spatial Selection for Efficient Cross-Modal Alignment in Referring Remote Sensing Image Segmentation

DGX agent

arXiv:2606.30244v1 Announce Type: new Abstract: Referring Remote Sensing Image Segmentation (RRSIS) seeks to localize and segment the target object or region specified by a natural language expression

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Sequential Hiring of Contingent Workers Through Learning-Based Optimization

DGX agent

arXiv:2606.18438v2 Announce Type: replace-cross Abstract: In this paper, we study a sequential workforce management problem in a contingent labor setting with uncertainty in both worker production and

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Sequential Planning via Anchored Robotic Keypoints

DGX agent

arXiv:2606.30613v1 Announce Type: new Abstract: We present Sequential Planning via Anchored Robotic Keypoints, SPARK, a training-free neurosymbolic manipulation system that reaches 43.7% on six LIBERO

model-releasesarxiv-cs-ro
30 Jun 2026
Model Releases

SEVA: Self-Evolving Verification Agent with Process Reward for Fact Attribution

DGX agent

arXiv:2606.29713v1 Announce Type: cross Abstract: Hallucination is the reliability bottleneck for LLM-based agents, and fact attribution verifiers are the last line of defense -- yet today's verifiers

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SFBench: The SciFy Scientific Feasibility Benchmark

DGX agent

arXiv:2606.29630v1 Announce Type: new Abstract: We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in material

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

DGX agent

arXiv:2606.30201v1 Announce Type: cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Singular Learning and Occam's Razor in Deep Monomial Networks

DGX agent

arXiv:2606.28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture. These critical poi

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

SkelEM: Training-Signal Decoupling of Skeleton and Diffusion for Self-supervised Axial Super-Resolution in Volume Microscopy

DGX agent

arXiv:2606.30012v1 Announce Type: new Abstract: Volume microscopy, including electron and light microscopy, suffers from severe anisotropic resolution due to physical axial sectioning. Existing self-s

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

Spatial Deconfounder: Interference-Aware Deconfounding for Spatial Causal Inference

DGX agent

arXiv:2510.08762v2 Announce Type: replace Abstract: Causal inference in spatial domains faces two intertwined challenges: (1) unmeasured spatial factors, such as weather, air pollution, or mobility, t

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Spectral Embedding via Chebyshev Bases for Robust DeepONet Approximation

DGX agent

arXiv:2512.09165v2 Announce Type: replace Abstract: Deep Operator Networks (DeepONets) have emerged as a powerful framework for data-driven operator learning, providing flexible surrogates for nonline

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

SpreadsheetBench 2: Evaluating Agents on End-to-End Business Spreadsheet Workflows

DGX agent

arXiv:2606.29955v1 Announce Type: cross Abstract: Spreadsheets are widely used for business analysis, financial modeling, reporting, and decision-making. However, most existing spreadsheet benchmarks

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

DGX agent

arXiv:2606.29815v1 Announce Type: new Abstract: Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by e

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Stability and Concentration in Nonlinear Inverse Problems with Block-Structured Parameters: Lipschitz Geometry, Identifiability, and an Application to Gaussian Splatting

DGX agent

arXiv:2602.09415v2 Announce Type: replace Abstract: We develop an operator-theoretic framework for stability and statistical concentration in nonlinear inverse problems with block-structured parameter

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley

DGX agent

arXiv:2507.07445v3 Announce Type: replace Abstract: Autonomous agents navigating human society must master both production activities and social interactions, yet existing benchmarks rarely evaluate t

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Statistically Indistinguishable, Operationally Distinct: A Formal Barrier for Tabular Foundation Models

DGX agent

arXiv:2606.29091v1 Announce Type: cross Abstract: Tabular foundation models cannot reason about data produced by running systems without access to the rules that govern them. We make this statement fa

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Stay Unique, Stay Efficient: Preserving Model Personality in Multi-Task Merging

DGX agent

arXiv:2512.01461v2 Announce Type: replace-cross Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training. However, traditional basic

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

STEMGym: Benchmarking Sequential Decision-Making under Dose Budgets in Autonomous Electron Microscopy

DGX agent

arXiv:2606.29592v1 Announce Type: new Abstract: A central premise of autonomous scientific imaging is that smarter navigation, whether Bayesian, RL-based, or otherwise adaptive, is the principal lever

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

StrucTab: A Structured Optimization Framework for Table Parsing

DGX agent

arXiv:2606.29905v1 Announce Type: new Abstract: Table parsing aims to convert table images into structured, machine-readable representations, a task requiring the joint perception of complex spatial l

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

SurgVLA-Bench: Towards Evaluating Vision-Language-Action Models for Laparoscopic Surgical Robotics

DGX agent

arXiv:2606.29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics. Despite the prevalence of VLA benchm

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SVC-Probe: A Framework for Evaluating Perturbation Generalization in Spatial Foundation-Model Embeddings

DGX agent

arXiv:2606.28465v1 Announce Type: cross Abstract: This work examines perturbation generalization in spatial foundation-model embeddings derived from fluorescence microscopy images. Although these mode

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SVCBench: A Streaming Video Counting Benchmark for Spatial-Temporal State Maintenance

DGX agent

arXiv:2603.12703v3 Announce Type: replace Abstract: Video understanding requires models to continuously track and update world state during playback. Although existing benchmarks have advanced video u

model-releasesarxiv-cs-cv
30 Jun 2026
Model Releases

SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?

DGX agent

arXiv:2511.06090v3 Announce Type: replace-cross Abstract: Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduce r

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SWE-Together: Evaluating Coding Agents in Interactive User Sessions

DGX agent

arXiv:2606.29957v1 Announce Type: cross Abstract: Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assi

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

SWITCH: Benchmarking Modeling and Handling of Tangible Interfaces in Long-horizon Embodied Scenarios

DGX agent

arXiv:2511.17649v4 Announce Type: replace-cross Abstract: Tangible control interfaces (TCIs), such as appliance panels, remotes, elevators, and embedded GUIs, are a fundamental component of everyday h

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Symbolic Mechanistic Data Attribution: Tracing Training Influence to Learned Behavioral Policies

DGX agent

arXiv:2606.29171v1 Announce Type: cross Abstract: While existing data attribution methods can identify which training examples build specific mechanistic circuits, they cannot explain how training dat

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

TextClusterLab: An Integrated Framework for Reliable Text Clustering Studies

DGX agent

arXiv:2606.28328v1 Announce Type: cross Abstract: In recent years, text clustering has become a critical technique for applications including intent discovery, topic mining, and recommendation systems

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

TF-MoE: Time-Frequency Mixture-of-Experts for Efficient Speech Separation

DGX agent

arXiv:2606.29575v1 Announce Type: cross Abstract: Recent advances in speech separation (SS) have led to compact front-end models with small parameter sizes, yet their high computational cost remains a

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The Complexity Ceiling Benchmark: A Multi-Domain Evaluation of Sequential Reasoning Under Depth Scaling

DGX agent

arXiv:2606.29278v1 Announce Type: new Abstract: We introduce the Complexity Ceiling Benchmark (CCB), a controlled evaluation of how language-model reasoning decays as the number of required sequential

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The Contagion Tensor: A Framework for Measuring Output-Distribution Coupling in Multi-Agent LLM Systems -- and Auditing the Claims It Enables

DGX agent

arXiv:2606.28839v1 Announce Type: new Abstract: We introduce the Contagion Tensor, a measurement framework for quantifying how large language model (LLM) output distributions couple across modalities,

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

The CRISTAL Method: Neurosymbolic analysis from AI-synthesized world models

DGX agent

arXiv:2606.29799v1 Announce Type: new Abstract: This project introduces the CRISTAL Method (Coherent Reliable Intentional Synthesis of Truthful Analysis Logic), a neurosymbolic framework for automatin

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

DGX agent

arXiv:2606.28325v1 Announce Type: cross Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a se

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

The FIL Hypothesis: Inductive Biases Help with Kernel Engineering

DGX agent

arXiv:2606.30442v1 Announce Type: new Abstract: The Bitter Lesson, which posits that general-purpose methods that scale with computation and data ultimately outperform those with built-in human knowle

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The Heterogeneous Safety Impacts of Benign Multilingual Fine-Tuning

DGX agent

arXiv:2606.28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task. However, prior work has shown th

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The Hidden Cost of Structured Generation in LLMs: Draft-Conditioned Constrained Decoding

DGX agent

arXiv:2603.03305v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used to generate executable outputs, JSON objects, and API calls, where a single syntax error ca

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The Human Creativity Benchmark

DGX agent

arXiv:2606.30561v1 Announce Type: new Abstract: Modern AI evaluation frameworks treat evaluator disagreement as noise to be resolved. In creative domains, professional disagreement reflects genuine di

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The Interference Gap: Comparing Retrieval Bounds in Human Memory and RAG Systems

DGX agent

arXiv:2606.28327v1 Announce Type: cross Abstract: How do retrieval bounds compare between human episodic memory and Retrieval-Augmented Generation (RAG) systems under semantic interference? We present

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

The NTNU System at the S&I Challenge 2025 SLA Open Track

DGX agent

arXiv:2506.05121v3 Announce Type: replace Abstract: A recent line of research on spoken language assessment (SLA) employs neural models such as BERT and wav2vec 2.0 (W2V) to evaluate speaking proficie

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

The Verbose Context Problem in Medical Records

DGX agent

arXiv:2606.29503v1 Announce Type: cross Abstract: The verbose context problem occurs when structured concepts have token-inefficient textual representations. This bottleneck is acute in population hea

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding

DGX agent

arXiv:2601.04693v2 Announce Type: replace Abstract: Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scar

model-releasesarxiv-cs-cl
30 Jun 2026
Model Releases

Toward an Energy-Optimized Operation of Data Centers Located in Wind Farms Using Reinforcement Learning

DGX agent

arXiv:2606.30316v1 Announce Type: new Abstract: This paper studies Reinforcement Learning as an online controller for curtailment-aware workload shifting in wind-turbine-integrated high-performance co

model-releasesarxiv-cs-lg
30 Jun 2026
Model Releases

Toward Secure and Reliable PDDL Formalization of Large Language Models with Planner-in-the-Loop Feedback

DGX agent

arXiv:2606.29700v1 Announce Type: new Abstract: Planning often requires symbolic specifications that are both executable and verifiable. For large language models deployed in autonomous or decision-su

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Towards Continual Motion-Language Agents: LoRA Variants for Incremental Motion Understanding and Generation

DGX agent

arXiv:2606.30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Towards Generalizable and Evidential Nuclear Magnetic Resonance-Based Molecular Structure Elucidation via Large Language Model Agent

DGX agent

arXiv:2606.29776v1 Announce Type: cross Abstract: Nuclear Magnetic Resonance (NMR) spectroscopy is the gold standard for molecular structure elucidation, yet interpreting complex spectra for unknown m

model-releasesarxiv-cs-ai
30 Jun 2026
Model Releases

Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation

DGX agent

arXiv:2606.30598v1 Announce Type: new Abstract: Estimating accurate 3D hand-object pose from in-the-wild egocentric RGB remains challenging due to severe occlusions and ambiguous contact. Existing lea

model-releasesarxiv-cs-cv
30 Jun 2026
← Previous
1…109110111112113…361
Next →