AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

S2O: Early Stopping for Sparse Attention via Online Permutation

DGX agent

arXiv:2602.22575v2 Announce Type: replace Abstract: Attention scales quadratically with sequence length, fundamentally limiting long-context inference. Existing block-granularity sparsification can re

model-releasesarxiv-cs-lg
6 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Safety and accuracy follow different scaling laws in clinical large language models

DGX agent

arXiv:2605.04039v1 Announce Type: new Abstract: Clinical LLMs are often scaled by increasing model size, context length, retrieval complexity, or inference-time compute, with the implicit expectation

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

SAM-NER: Semantic Archetype Mediation for Zero-Shot Named Entity Recognition

DGX agent

arXiv:2605.03706v1 Announce Type: new Abstract: Zero-shot Named Entity Recognition (ZS-NER) remains brittle under domain and schema shifts, where unseen label definitions often misalign with a large l

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Scaling Unsupervised Multi-Source Federated Domain Adaptation through Group-Wise Discrepancy Minimization

DGX agent

arXiv:2510.08150v3 Announce Type: replace Abstract: Unsupervised multi-source domain adaptation (UMDA) leverages labeled data from multiple source domains to generalize to an unlabeled target. While f

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

SCGNN: Semantic Consistency enhanced Graph Neural Network Guided by Granular-ball Computing

DGX agent

arXiv:2605.02617v2 Announce Type: new Abstract: Capturing semantic consistency among nodes is crucial for effective graph representation learning. Existing approaches typically rely on k-nearest neigh

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Seeking Information with RAG-Assistants: Does Model Size Matter in Human-AI Collaborations?

DGX agent

arXiv:2605.00964v1 Announce Type: cross Abstract: Much research on LLMs has focused on increasing benchmark performance. However, the evaluation of such models in real-world collaborative human-AI wor

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Self-Mined Hardness for Safety Fine-Tuning

DGX agent

arXiv:2605.03226v1 Announce Type: new Abstract: Safety fine-tuning of language models typically requires a curated adversarial dataset. We take a different approach: score each candidate prompt's diff

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning

DGX agent

arXiv:2605.03189v1 Announce Type: new Abstract: Image captioning has become an important task in computer vision, enabling models to generate natural language descriptions of visual content. While sev

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Simulated Students in Tutoring Dialogues: Substance or Illusion?

DGX agent

arXiv:2601.04025v2 Announce Type: replace Abstract: Advances in large language models (LLMs) enable many new innovations in education. However, evaluating the effectiveness of new technology requires

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Soft Tournament Equilibrium

DGX agent

arXiv:2604.04328v3 Announce Type: replace-cross Abstract: The evaluation of general-purpose artificial agents, particularly those based on LLMs, presents a significant challenge due to the non-transit

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Sparse Memory Finetuning as a Low-Forgetting Alternative to LoRA and Full Finetuning

DGX agent

arXiv:2605.03229v1 Announce Type: new Abstract: Adapting a pretrained language model to a new task often hurts the general capabilities it already had, a problem known as catastrophic forgetting. Spar

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Stable Multimodal Graph Unlearning via Feature-Dimension Aware Quantile Selection

DGX agent

arXiv:2605.03303v1 Announce Type: new Abstract: Graph unlearning remains a critical technique for supporting privacy-preserving and sustainable multimodal graph learning. However, we observe that exis

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

StateSMix: Online Lossless Compression via Mamba State Space Models and Sparse N-gram Context Mixing

DGX agent

arXiv:2605.02904v1 Announce Type: new Abstract: We present StateSMix, a fully self-contained lossless compressor that couples an online-trained Mamba-style State Space Model (SSM) with sparse n-gram c

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

StateVLM: A State-Aware Vision-Language Model for Robotic Affordance Reasoning

DGX agent

arXiv:2605.03927v1 Announce Type: new Abstract: Vision-language models (VLMs) have shown remarkable performance in various robotic tasks, as they can perceive visual information and understand natural

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

SURE-RAG: Sufficiency and Uncertainty-Aware Evidence Verification for Selective Retrieval-Augmented Generation

DGX agent

arXiv:2605.03534v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, but retrieval is not verification: a passage can be topical and still fail t

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Task-Aware Scanning Parameter Configuration for Robotic Inspection Using Vision Language Embeddings and Hyperdimensional Computing

DGX agent

arXiv:2605.03909v1 Announce Type: cross Abstract: Robotic laser profiling is widely used for dimensional verification and surface inspection, yet measurement fidelity is often dominated by sensor conf

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Tenability and Weak Semantics: Modeling Non-uniform Defense -- Extended Version

DGX agent

arXiv:2605.02024v1 Announce Type: new Abstract: In Dung-style abstract argumentation, various semantics capture notions of acceptability of arguments. The admissibility semantics capture the notion th

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

The Cost of Consensus: Isolated Self-Correction Prevails Over Unguided Homogeneous Multi-Agent Debate

DGX agent

arXiv:2605.00914v1 Announce Type: cross Abstract: Multi-agent debate, where teams of LLMs iteratively exchange rationales and vote on answers, is widely deployed under the assumption that peer review

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

The Dynamic Gist-Based Memory Model (DGMM): A Memory-Centric Architecture for Artificial Intelligence

DGX agent

arXiv:2605.02106v1 Announce Type: new Abstract: Contemporary artificial intelligence systems achieve strong performance through large-scale parameterization, retrieval augmentation, and training on ex

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

The Oracle's Fingerprint: Correlated AI Forecasting Errors and the Limits of Bias Transmission

DGX agent

arXiv:2605.00844v1 Announce Type: cross Abstract: When large language models (LLMs) are consulted as forecasting tools, the independence of individual errors -- the foundation of collective intelligen

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

The Productivity-Reliability Paradox: Specification-Driven Governance for AI-Augmented Software Development

DGX agent

arXiv:2605.01160v1 Announce Type: cross Abstract: Since 2022, AI-powered coding assistants have produced contradictory evidence: controlled studies report 20-56% productivity gains on well-scoped task

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It

DGX agent

arXiv:2605.03258v1 Announce Type: cross Abstract: Large language models often fail at simple counting tasks, even when the items to count are explicitly present in the prompt. We investigate whether t

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

The TTS-STT Flywheel: Synthetic Entity-Dense Audio Closes the Indic ASR Gap Where Commercial and Open-Source Systems Fail

DGX agent

arXiv:2605.03073v1 Announce Type: new Abstract: Niche-domain Indic ASR -- digit strings, currency amounts, addresses, brand names, English/Indic codemix -- is under-served by both open-source SOTA and

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

TMD-Bench: A Multi-Level Evaluation Paradigm for Music-Dance Co-Generation

DGX agent

arXiv:2605.01809v1 Announce Type: cross Abstract: Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive medi

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Toward a Science of Intent: Closure Gaps and Delegation Envelopes for Open-World AI Agents

DGX agent

arXiv:2604.25000v2 Announce Type: replace Abstract: Recent work has framed intelligence in verifiable tasks as reducing time-to-solution through learned structure and test-time search, while systems w

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Toward Generative Quantum Utility via Correlation-Complexity Map

DGX agent

arXiv:2603.06440v2 Announce Type: replace Abstract: We study a practical question in generative quantum machine learning: given a classical dataset, can we determine, before training, whether it is we

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

Towards Agentic Runtime Healing

DGX agent

arXiv:2408.01055v2 Announce Type: replace-cross Abstract: Self-healing systems have long been a focus of research, aiming to enable software to recover from unexpected runtime errors without human int

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Towards Multi-Agent Autonomous Reasoning in Hydrodynamics

DGX agent

arXiv:2605.01102v1 Announce Type: new Abstract: Single-agent systems (SAS) have become the default pattern for LLM-driven scientific workflows, but routing planning, tool use, and synthesis through a

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Towards Understanding Specification Gaming in Reasoning Models

DGX agent

arXiv:2605.02269v1 Announce Type: new Abstract: Specification gaming is a critical failure mode of LLM agents. Despite this, there has been little systematic research into when it arises and what driv

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Training-Free Probabilistic Time-Series Forecasting with Conformal Seasonal Pools

DGX agent

arXiv:2605.03789v1 Announce Type: cross Abstract: We propose Conformal Seasonal Pools (CSP), a training-free probabilistic time-series forecaster that mixes same-season empirical draws with signed res

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

TriBench-Ko: Evaluating LLM Risks in Judicial Workflows

DGX agent

arXiv:2605.03792v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into legal workflows. However, existing benchmarks primarily address proxy tasks, such as bar e

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

Trojan Hippo: Weaponizing Agent Memory for Data Exfiltration

DGX agent

arXiv:2605.01970v2 Announce Type: cross Abstract: Memory systems enable otherwise-stateless LLM agents to persist user information across sessions, but also introduce a new attack surface. We characte

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

UnAC: Adaptive Visual Prompting with Abstraction and Stepwise Checking for Complex Multimodal Reasoning

DGX agent

arXiv:2605.03950v1 Announce Type: new Abstract: Although recent LMMs have become much stronger at visual perception, they remain unreliable on problems that require multi-step reasoning over visual ev

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Valley3: Scaling Omni Foundation Models for E-commerce

DGX agent

arXiv:2605.01278v1 Announce Type: new Abstract: In this work, we present Valley3, an omni multimodal large language model (MLLM) developed for diverse global e-commerce tasks, with unified understandi

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Vanishing L2 regularization for the softmax Multi Armed Bandit

DGX agent

arXiv:2605.03752v1 Announce Type: new Abstract: Multi Armed Bandit (MAB) algorithms are a cornerstone of reinforcement learning and have been studied both theoretically and numerically. One of the mos

model-releasesarxiv-cs-lg
6 May 2026
Model Releases

VEBench:Benchmarking Large Multimodal Models for Real-World Video Editing

DGX agent

arXiv:2605.03276v1 Announce Type: new Abstract: Real-world video editing demands not only expert knowledge of cinematic techniques but also multimodal reasoning to select, align, and combine footage i

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development

DGX agent

arXiv:2603.04601v2 Announce Type: replace-cross Abstract: Code generation has emerged as one of AI's highest-impact use cases, yet existing benchmarks measure isolated tasks rather than the complete '

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

VisInject: Disruption != Injection -- A Dual-Dimension Evaluation of Universal Adversarial Attacks on Vision-Language Models

DGX agent

arXiv:2605.01449v1 Announce Type: cross Abstract: Universal adversarial attacks on aligned multimodal large language models are increasingly reported with attack success rates in the 60-80% range, sug

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models

DGX agent

arXiv:2605.03351v1 Announce Type: new Abstract: Video vision-language models (VLMs) keep paying for visual state the stream already told us was stable. The factory wall did not move, but most VLM pipe

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

What Makes VLMs Robust? Towards Reconciling Robustness and Accuracy in Vision-Language Models

DGX agent

arXiv:2603.12799v2 Announce Type: replace Abstract: Achieving adversarial robustness in Vision-Language Models (VLMs) inevitably compromises accuracy on clean data, presenting a long-standing and chal

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

When Alignment Isn't Enough: Response-Path Attacks on LLM Agents

DGX agent

arXiv:2605.02187v1 Announce Type: cross Abstract: Bring-Your-Own-Key (BYOK) agent architectures let users route LLM traffic through third-party relays, creating a critical integrity gap: a malicious r

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

When Prompts Interact: Assessing Prompt Arithmetic for Deconfounding under Distribution Shift

DGX agent

arXiv:2605.03096v1 Announce Type: cross Abstract: In classification tasks, models may rely on confounding variables to achieve strong in-distribution performance, capturing spurious features that fail

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

When Should a Language Model Trust Itself? Same-Model Self-Verification as a Conditional Confidence Signal

DGX agent

arXiv:2605.02915v1 Announce Type: new Abstract: Same-model self-verification, prompting a model to audit its own predicted answer, is a plausible confidence signal for selective prediction, but its pr

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

When Stress Becomes Signal: Detecting Antifragility-Compatible Regimes in Multi-Agent LLM Systems

DGX agent

arXiv:2605.02463v1 Announce Type: cross Abstract: Multi-agent LLM systems are increasingly used to solve complex tasks through decomposition, debate, specialization, and ensemble reasoning. However, t

model-releasesarxiv-cs-ai
6 May 2026
Model Releases

Where to Bind Matters: Hebbian Fast Weights in Vision Transformers for Few-Shot Character Recognition

DGX agent

arXiv:2605.02920v1 Announce Type: cross Abstract: Standard transformer architectures learn fixed slow-weight representations during training and lack mechanisms for rapid adaptation within an episode.

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File Dependencies

DGX agent

arXiv:2605.03596v1 Announce Type: cross Abstract: Workspace learning requires AI agents to identify, reason over, exploit, and update explicit and implicit dependencies among heterogeneous files in a

model-releasesarxiv-cs-cl
6 May 2026
Model Releases

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models

DGX agent

arXiv:2605.03475v1 Announce Type: new Abstract: Evaluating generative video models remains an open problem. Reference-based metrics such as Structural Similarity Index Measure (SSIM) and Peak Signal t

model-releasesarxiv-cs-cv
6 May 2026
Model Releases

A Closed-Form Persistence-Landmark Pipeline for Certified Point-Cloud and Graph Classification

DGX agent

arXiv:2605.02836v1 Announce Type: new Abstract: We introduce PLACE (Persistence-Landmark Analytic Classification Engine), a closed-form pipeline for classifying point clouds and graphs through their p

model-releasesarxiv-cs-lg
5 May 2026
← Previous
1…275276277278279…361
Next →