AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

Metadata Predictability Is Not Evidence Dependence: An Intervention-Based Audit for Weak-Label Benchmarks

DGX agent

arXiv:2605.23701v1 Announce Type: new Abstract: We study a protocol-level test for weak-label benchmarks: whether benchmark outputs change when the provided evidence is intervened on. Metadata-only sh

model-releasesarxiv-cs-cl
25 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Model Collapse as Cultural Evolution

DGX agent

arXiv:2605.23054v1 Announce Type: cross Abstract: Model collapse, the progressive degradation of LLMs trained on their own outputs, has been characterized statistically but lacks a linguistic explanat

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU

DGX agent

arXiv:2605.23057v1 Announce Type: cross Abstract: ModeSwitch-LLM is a lightweight request-boundary controller for improving single-GPU large language model inference efficiency by routing each request

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Moonwalk: Inverse-Forward Differentiation

DGX agent

arXiv:2402.14212v4 Announce Type: replace-cross Abstract: Backpropagation's main limitation is its need to store intermediate activations (residuals) during the forward pass, which restricts the depth

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer

DGX agent

arXiv:2605.23871v1 Announce Type: cross Abstract: We develop a gradient flow on the space of probability measures defined on matrix-valued parameters induced by regularized Muon, an analytically smoot

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models

DGX agent

arXiv:2505.17015v2 Announce Type: replace-cross Abstract: Multi-modal large language models (MLLMs) have rapidly advanced in visual tasks, yet their spatial understanding remains limited to single ima

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection

DGX agent

arXiv:2605.23036v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in large language models (LLMs), but SAE-based lang

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Non-normal spectral signatures of instability in neural network training dynamics

DGX agent

arXiv:2605.23476v1 Announce Type: new Abstract: Training instabilities in deep networks - loss spikes, oscillatory convergence, and gradient pathologies - are empirically prevalent but lack a rigorous

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

NP-LoRA: Null Space Projection for Subject-Style LoRA Fusion

DGX agent

arXiv:2511.11051v3 Announce Type: replace Abstract: Low-Rank Adaptation (LoRA) fusion enables the composition of subject and style representations for controllable generation without retraining. Howev

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

One Policy, Infinite NPCs: Persona-Traceable Shared RL Policies for Scalable Game Agents

DGX agent

arXiv:2605.23652v1 Announce Type: new Abstract: On a 300-persona life-simulation benchmark, pcsp achieves compositional zero-shot persona identification up to 17x above chance, Spearman rho approx 0.7

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Online Partitioned Local Depth for semi-supervised applications

DGX agent

arXiv:2512.15436v2 Announce Type: replace-cross Abstract: We introduce an extension of the partitioned local depth (PaLD) algorithm that is adapted to online applications such as semi-supervised predi

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Ontological Knowledge Blocks: Executable Compliance and Profile-Based Validation for Trustworthy AI Systems

DGX agent

arXiv:2605.23297v1 Announce Type: new Abstract: AI-enabled services deployed in critical digital infrastructure are subject to governance obligations spanning transparency, accountability, fairness, a

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Open Multimodal Datasets and Open-Source Software for Data-Driven Modeling of Multiphase Transport and Thermal Systems

DGX agent

arXiv:2605.23037v1 Announce Type: new Abstract: Data-driven modeling is becoming central to multiphase transport, electronics cooling, acoustic diagnostics, and thermal-fluid digital twins, but progre

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

DGX agent

arXiv:2605.23657v1 Announce Type: new Abstract: Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improvin

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

Operator Learning for Reconstructing Flow Fields from Sparse Measurements: a Language Model Approach

DGX agent

arXiv:2605.23712v1 Announce Type: cross Abstract: Reconstructing flow fields from sparse measurements is a fundamental problem in fluid mechanics with broad implications for modeling, control, and des

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Optimization of randomized neural networks for transfer operator approximation

DGX agent

arXiv:2605.23689v1 Announce Type: new Abstract: RaNNDy is a randomized neural network architecture for the data-driven approximation of transfer operators associated with complex dynamical systems. Th

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Order-Optimal Sequential 1-Bit Mean Estimation in General Tail Regimes

DGX agent

arXiv:2604.07796v2 Announce Type: replace-cross Abstract: In this paper, we study the problem of mean estimation under 1-bit communication constraints. We propose a novel adaptive mean estimator based

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

PACE: Two-Timescale Self-Evolution for Small Language Model Agents

DGX agent

arXiv:2605.23019v1 Announce Type: new Abstract: Deploying language-model agents in production often requires substantial compute and human effort to tune prompts, parsers, validators, and other compon

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Parallel Context Compaction for Long-Horizon LLM Agent Serving

DGX agent

arXiv:2605.23296v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate growing conversation histories that eventually exceed the model's context window. Context compaction via LLM-based su

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

PGT: Procedurally Generated Tasks for improving visual grounding in MLLMs

DGX agent

arXiv:2605.23883v1 Announce Type: cross Abstract: Despite remarkable progress in Multimodal Large Language Models (MLLMs), these models still struggle with fine-grained understanding tasks. In this wo

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Philosophical Dispositions as Behavioral Constraints for AI-Assisted Code Review: An Empirical Study

DGX agent

arXiv:2605.23108v1 Announce Type: cross Abstract: AI-assisted code review tools typically operate as generic 'expert reviewer' agents, producing homogeneous findings regardless of the analysis type ne

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

PhotoFlow: Agentic 3D Virtual Photography Missions

DGX agent

arXiv:2605.23771v1 Announce Type: cross Abstract: Virtual photography asks an agent to enter a prepared 3D scene with no preselected camera pose or reference image, infer a suitable shot from scene in

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Physics-Informed Machine Learning Regulated by Finite Element Analysis for Simulation Acceleration of Melt Pool Dynamics in Laser Powder Bed Fusion

DGX agent

arXiv:2506.20537v3 Announce Type: replace Abstract: Efficient simulation of Laser Powder Bed Fusion (LPBF) is crucial for process prediction due to the lasting issue of high computational cost associa

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Physiome-ODE: A Benchmark for Irregularly Sampled Multivariate Time Series Forecasting Based on Biological ODEs

DGX agent

arXiv:2502.07489v2 Announce Type: replace Abstract: State-of-the-art methods for forecasting irregularly sampled time series with missing values predominantly rely on just four datasets and a few smal

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

PixelPonder: Dynamic Patch Adaptation for Enhanced Multi-Conditional Text-to-Image Generation

DGX agent

arXiv:2503.06684v3 Announce Type: replace Abstract: Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Pointwise Metrics Mislead: An Evaluation Protocol for Multimodal Inverse Problems

DGX agent

arXiv:2605.22891v1 Announce Type: new Abstract: Evaluation in scientific reconstruction is dominated by pointwise metrics - RMSE, MAE, per-event resolution - under the implicit assumption that lower e

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

PoisonForge: Task-Level Targeted Poisoning Benchmark for Instruction-Tuned LLMs

DGX agent

arXiv:2605.23168v1 Announce Type: cross Abstract: When practitioners fine-tune LLMs on unvetted datasets, an adversary can exploit the data supply chain through task-level poisoning: inserting a small

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks

DGX agent

arXiv:2605.23170v1 Announce Type: cross Abstract: Position-controlled evaluation is standard for retrieval tasks such as Needle-in-a-Haystack and RULER, but mainstream reasoning benchmarks do not cont

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations

DGX agent

arXiv:2605.22855v1 Announce Type: cross Abstract: Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision makin

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

PROGRESSLM: Towards Progress Reasoning in Vision-Language Models

DGX agent

arXiv:2601.15224v2 Announce Type: replace-cross Abstract: Estimating task progress requires reasoning over long-horizon dynamics rather than recognizing static visual content. While modern Vision-Lang

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

ProtDBench: A Unified Benchmark of Protein Binder Design and Evaluation

DGX agent

arXiv:2605.04118v2 Announce Type: replace-cross Abstract: Recent advances in de novo protein binder design have enabled increasing experimental validation, yet reported in silico metrics remain diffic

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents

DGX agent

arXiv:2605.23574v1 Announce Type: new Abstract: Long-horizon language agents can make many plausible local tool calls yet fail to persist until a requested count is actually complete. We study this ga

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

R^3L: Reflect-then-Retry Reinforcement Learning with Language-Guided Exploration, Pivotal Credit, and Positive Amplification

DGX agent

arXiv:2601.03715v2 Announce Type: replace-cross Abstract: Reinforcement learning drives recent advances in LLM reasoning and agentic capabilities, yet current approaches struggle with both exploration

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Recursive Block-Diagonal Coupling for Resource-Efficient Training of Vision Models

DGX agent

arXiv:2605.23656v1 Announce Type: new Abstract: Training high-capacity vision models from scratch requires substantial computational resources. To improve training efficiency of a wide target model, e

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Reinforcement Learning for Microcanonical Graph Ensemble with Assortativity Constraints

DGX agent

arXiv:2605.23285v1 Announce Type: cross Abstract: How network structure determines function is a fundamental question, and it can be investigated by graph ensembles with precisely controlled structura

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Resilience Characterization of AI-Native Wireless Receivers via Persistent Homology

DGX agent

arXiv:2605.22886v1 Announce Type: cross Abstract: AI-native wireless receivers based on deep learning exhibit remarkable performance under stationary channel conditions, yet their resilience to distri

model-releasesarxiv-cs-lg
25 May 2026
Model Releases

Revitalizing Dense Material Segmentation: Stabilized Vision Transformers and the Generalization Paradox

DGX agent

arXiv:2605.23747v1 Announce Type: new Abstract: Material segmentation, the pixel-wise classification of physical surface properties, remains a challenging problem in computer vision, requiring physico

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

RMA: an Agentic System for Research-Level Mathematical Problems

DGX agent

arXiv:2605.22875v1 Announce Type: new Abstract: We present extbf{Research Math Agents (RMA)}, an agentic framework for automated reasoning on research-level mathematical problems. Unlike prior studies

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

RoboSurg-VQA: A Multimodal Benchmark for Surgical Segmentation-Aware Visual Question Answering

DGX agent

arXiv:2605.23068v1 Announce Type: new Abstract: Reliable visual understanding in robot-assisted and minimally invasive surgery (RMIS/MIS) demands more than accurate masks: in clinical practice, clinic

model-releasesarxiv-cs-cv
25 May 2026
Model Releases

Same Model, Different Weakness: How Language and Modality Reshape the Jailbreak Attack Surface in Frontier MLLMs

DGX agent

arXiv:2605.23157v1 Announce Type: new Abstract: The attack surface of a multimodal large language model (MLLM) is language-dependent in ways that reveal the mechanistic structure of alignment failures

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

SciAtlas: A Large-Scale Knowledge Graph for Automated Scientific Research

DGX agent

arXiv:2605.22878v1 Announce Type: new Abstract: The exponential growth of global academic output has confronted researchers and AI agents with an unprecedented ``information explosion,'' where fragmen

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

SciHorizon-GENE: Benchmarking LLM for Life Sciences Inference from Gene Knowledge to Functional Understanding

DGX agent

arXiv:2601.12805v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown growing promise in biomedical research, particularly for knowledge-driven interpretation tasks. Howeve

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

DGX agent

arXiv:2605.22903v1 Announce Type: cross Abstract: Benchmark accuracy is often implicitly assumed to reflect grounded visual understanding in vision-language models (VLMs), yet it remains unclear to wh

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Semantically Structured Mixture-of-Experts for Compositional Robotic Manipulation

DGX agent

arXiv:2605.23477v1 Announce Type: new Abstract: Diffusion-based policies have established a new standard for precise robotic manipulation but face a critical scalability bottleneck: high-performance m

model-releasesarxiv-cs-ro
25 May 2026
Model Releases

SemEval-2026 Task 6: CLARITY -- Unmasking Political Question Evasions

DGX agent

arXiv:2603.14027v2 Announce Type: replace Abstract: Political speakers often avoid answering questions directly while maintaining the appearance of responsiveness. Despite its importance for public di

model-releasesarxiv-cs-cl
25 May 2026
Model Releases

SkillOpt: Executive Strategy for Self-Evolving Agent Skills

DGX agent

arXiv:2605.23904v1 Announce Type: new Abstract: Agent skills today are hand-crafted, generated one-shot, or evolved through loosely controlled self-revision, none of which behaves like a deep-learning

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Sparse Autoencoders Map Brain-LLM Alignment onto Cortical Semantic Topography

DGX agent

arXiv:2605.23035v1 Announce Type: cross Abstract: Intermediate layers of large language models (LLMs) best predict human brain responses to language, one of the most robust findings in computational n

model-releasesarxiv-cs-ai
25 May 2026
Model Releases

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation

DGX agent

arXiv:2412.14642v4 Announce Type: replace Abstract: Recently, Large Language Models (LLMs) have demonstrated great potential in natural language-driven molecule discovery. However, existing datasets a

model-releasesarxiv-cs-cl
25 May 2026
← Previous
1…206207208209210…361
Next →