AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,460
  • Agents7,259
  • Applications5,196
  • Concepts5
  • Hardware1,748
  • Industry6,091
  • Local Ai4,708
  • Model Releases22,512
  • Research19,191
  • Safety12,809
  • Syntheses17
  • Tools1,665
  • Tutorials3,259

Source
HumanDGX agent

Content type
84,460Total entries
1Added by human
84,459Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

AgenticRecTune: Multi-Agent with Self-Evolving Skillhub for Recommendation System Optimization

DGX agent

arXiv:2604.26969v1 Announce Type: cross Abstract: Modern large-scale recommendation systems are typically constructed as multi-stage pipelines, encompassing pre-ranking, ranking, and re-ranking phases

model-releasesarxiv-cs-ai
1 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

APPSI-139: A Parallel Corpus of English Application Privacy Policy Summarization and Interpretation

DGX agent

arXiv:2604.27550v1 Announce Type: cross Abstract: Privacy policies are essential for users to understand how service providers handle their personal data. However, these documents are often long and c

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR

DGX agent

arXiv:2604.27543v1 Announce Type: new Abstract: Evaluating English ASR systems for conversational AI applications remains difficult, as many publicly available corpora are either pre-segmented into sh

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Assessing Pancreatic Ductal Adenocarcinoma Vascular Invasion: the PDACVI Benchmark

DGX agent

arXiv:2604.27582v1 Announce Type: new Abstract: Surgical resection remains the only potentially curative treatment for pancreatic ductal adenocarcinoma (PDAC), and eligibility depends on accurate asse

model-releasesarxiv-cs-cv
1 May 2026
Model Releases

Auditing Frontier Vision-Language Models for Trustworthy Medical VQA: Grounding Failures, Format Collapse, and Domain Adaptation

DGX agent

arXiv:2604.27720v1 Announce Type: new Abstract: Deploying vision-language models (VLMs) in clinical settings demands auditable behavior under realistic failure conditions, yet the failure landscape of

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Auto-FlexSwitch: Efficient Dynamic Model Merging via Learnable Task Vector Compression

DGX agent

arXiv:2604.28109v1 Announce Type: new Abstract: Model merging has attracted attention as an effective path toward multi-task adaptation by integrating knowledge from multiple task-specific models. Amo

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Autoregressive Synthesis of Sparse and Semi-Structured Mixed-Type Data

DGX agent

arXiv:2603.01444v2 Announce Type: replace Abstract: Synthetic data generation is an important capability for privacy-preserving data sharing, system benchmarking and test data provisioning. For mixed-

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism

DGX agent

arXiv:2604.27089v1 Announce Type: new Abstract: Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of tho

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

AutoSurfer -- Teaching Web Agents through Comprehensive Surfing, Learning, and Modeling

DGX agent

arXiv:2604.27253v1 Announce Type: new Abstract: Recent advances in multimodal large language models (LLMs) have revolutionized web agents that can automate complex tasks on websites. However, their ac

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

BatteryPass-12K: The First Dataset for the Novel Digital Battery Passport Conformance Task

DGX agent

arXiv:2604.26986v1 Announce Type: new Abstract: We introduce a novel task of digital battery passport (DBP) conformance classification and introduce the first public benchmark for the task: BatteryPas

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Bayesian Hierarchical Models and the Maximum Entropy Principle

DGX agent

arXiv:2603.10252v2 Announce Type: replace-cross Abstract: Bayesian hierarchical models are frequently used in practical data analysis contexts. One interpretation of these models is that they provide

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Bayesian X-Learner: Calibrated Posterior Inference for Heterogeneous Treatment Effects under Heavy-Tailed Outcomes

DGX agent

arXiv:2604.27394v1 Announce Type: cross Abstract: Conditional Average Treatment Effect (CATE) estimation in practice demands three properties simultaneously: heterogeneous effects au(x), calibrated un

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Better Models, Faster Training: Sigmoid Attention for single-cell Foundation Models

DGX agent

arXiv:2604.27124v1 Announce Type: new Abstract: Training stable biological foundation models requires rethinking attention mechanisms: we find that using sigmoid attention as a drop in replacement for

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs

DGX agent

arXiv:2604.27006v1 Announce Type: cross Abstract: Context: Study screening in systematic literature reviews is costly, inconsistency-prone, and risk-asymmetric, since false negatives can compromise va

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Beyond Semantics: Measuring Fine-Grained Emotion Preservation in Small Language Model-Based Machine Translation

DGX agent

arXiv:2604.27920v1 Announce Type: cross Abstract: Preserving affective nuance remains a challenge in Machine Translation (MT), where semantic equivalence often takes precedence over emotional fidelity

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation

DGX agent

arXiv:2604.27405v1 Announce Type: cross Abstract: We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro i

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

BoostLoRA: Growing Effective Rank by Boosting Adapters

DGX agent

arXiv:2604.27308v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning (PEFT) methods face a tradeoff between adapter size and expressivity: ultra-low-parameter adapters are confined to fix

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Can Large Language Models Implement Agent-Based Models? An ODD-based Replication Study

DGX agent

arXiv:2602.10140v2 Announce Type: replace-cross Abstract: Large language models (LLMs) can now synthesize non-trivial executable code from textual descriptions, raising an important question: can LLMs

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

CareGuardAI: Context-Aware Multi-Agent Guardrails for Clinical Safety & Hallucination Mitigation in Patient-Facing LLMs

DGX agent

arXiv:2604.26959v1 Announce Type: cross Abstract: Integrating large language models (LLMs) into patient-facing healthcare systems offers significant potential to improve access to medical information.

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

CausalCompass: Evaluating the Robustness of Time-Series Causal Discovery in Misspecified Scenarios

DGX agent

arXiv:2602.07915v2 Announce Type: replace-cross Abstract: Causal discovery from time series is a fundamental task in machine learning. However, its widespread adoption is hindered by a reliance on unt

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Characterizing the Consistency of the Emergent Misalignment Persona

DGX agent

arXiv:2604.28082v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) on narrowly misaligned data generalizes to broadly misaligned behavior, a phenomenon termed emergent misalignme

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

ChipLingo: A Systematic Training Framework for Large Language Models in EDA

DGX agent

arXiv:2604.27415v1 Announce Type: new Abstract: With the rapid advancement of semiconductor technology, Electronic Design Automation (EDA) has become an increasingly knowledge-intensive and document-d

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

CL-bench Life: Can Language Models Learn from Real-Life Context?

DGX agent

arXiv:2604.27043v1 Announce Type: new Abstract: Today's AI assistants such as OpenClaw are designed to handle context effectively, making context learning an increasingly important capability for mode

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows

DGX agent

arXiv:2604.28139v1 Announce Type: cross Abstract: LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

COHERENCE: Benchmarking Fine-Grained Image-Text Alignment in Interleaved Multimodal Contexts

DGX agent

arXiv:2604.27389v1 Announce Type: cross Abstract: In recent years, Multimodal Large Language Models (MLLMs) have achieved remarkable progress on a wide range of multimodal benchmarks. Despite these ad

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Contextual Agentic Memory is a Memo, Not True Memory

DGX agent

arXiv:2604.27707v1 Announce Type: new Abstract: Current agentic memory systems (vector stores, retrieval-augmented generation, scratchpads, and context-window management) do not implement memory: they

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Cross-Lingual Response Consistency in Large Language Models: An ILR-Informed Evaluation of Claude Across Six Languages

DGX agent

arXiv:2604.27137v1 Announce Type: new Abstract: This paper introduces a systematic evaluation framework grounded in the Interagency Language Roundtable (ILR) Skill Level Descriptions and applies it to

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Decoding Scientific Experimental Images: The SPUR Benchmark for Perception, Understanding, and Reasoning

DGX agent

arXiv:2604.27604v1 Announce Type: new Abstract: We introduce SPUR, a comprehensive benchmark for scientific experimental image perception, understanding, and reasoning, comprising 4,264 question-answe

model-releasesarxiv-cs-cv
1 May 2026
Model Releases

DeepTutor: Towards Agentic Personalized Tutoring

DGX agent

arXiv:2604.26962v1 Announce Type: cross Abstract: Education represents one of the most promising real-world applications for Large Language Models (LLMs). However, conventional tutoring systems rely o

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

DEFault++: Automated Fault Detection, Categorization, and Diagnosis for Transformer Architectures

DGX agent

arXiv:2604.28118v1 Announce Type: cross Abstract: Transformer models are widely deployed in critical AI applications, yet faults in their attention mechanisms, projections, and other internal componen

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Diagnosing Capability Gaps in Fine-Tuning Data

DGX agent

arXiv:2604.27547v1 Announce Type: new Abstract: Fine-tuning large language models (LLMs) for domain-specific tasks requires training datasets that comprehensively cover the target capabilities a pract

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Do Papers Tell the Whole Story? A Benchmark and Framework for Uncovering Hidden Implementation Gaps in Bioinformatics

DGX agent

arXiv:2603.22018v2 Announce Type: replace Abstract: Ensuring consistency between research papers and their corresponding software code implementations is a fundamental prerequisite for guaranteeing th

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Do What I Say: A Spoken Prompt Dataset for Instruction-Following

DGX agent

arXiv:2603.09881v2 Announce Type: replace Abstract: Speech Large Language Models (SLLMs) have rapidly expanded, supporting a wide range of tasks. These models are typically evaluated using text prompt

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Do World Action Models Generalize Better than VLAs? A Robustness Study

DGX agent

arXiv:2603.22078v3 Announce Type: replace Abstract: Robot action planning in the real world is challenging as it requires not only understanding the current state of the environment but also predictin

model-releasesarxiv-cs-ro
1 May 2026
Model Releases

DPN-LE: Dual Personality Neuron Localization and Editing for Large Language Models

DGX agent

arXiv:2604.27929v1 Announce Type: new Abstract: With the widespread adoption of large language models (LLMs), understanding their personality representation mechanisms has become critical. As a novel

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Dynamic Scaled Gradient Descent for Stable Fine-Tuning for Classifications

DGX agent

arXiv:2604.27987v1 Announce Type: new Abstract: Fine-tuning pretrained models has become a standard approach to adapting pretrained knowledge to improve the accuracy on new sparse, imbalance datasets.

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

EdgeSpike: Spiking Neural Networks for Low-Power Autonomous Sensing in Edge IoT Architectures

DGX agent

arXiv:2604.27004v1 Announce Type: cross Abstract: We propose EdgeSpike, a co-designed spiking neural network (SNN) framework for autonomous low-power sensing in edge Internet of Things (IoT) architect

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

EDU-CIRCUIT-HW: Evaluating Multimodal Large Language Models on Real-World University-Level STEM Student Handwritten Solutions

DGX agent

arXiv:2602.00095v3 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) hold significant promise for revolutionizing traditional education and reducing teachers' workload. H

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Entropy of Ukrainian

DGX agent

arXiv:2604.27534v1 Announce Type: new Abstract: In natural language processing, the entropy of a language is a measure of its unpredictability and complexity. The first study on this subject was condu

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Exploring Interaction Paradigms for LLM Agents in Scientific Visualization

DGX agent

arXiv:2604.27996v1 Announce Type: new Abstract: This paper examines how different types of large language model (LLM) agents perform on scientific visualization (SciVis) tasks, where users generate vi

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

Exploring the Limits of Pruning: Task-Specific Neurons, Model Collapse, and Recovery in Task-Specific Large Language Models

DGX agent

arXiv:2604.27115v1 Announce Type: new Abstract: Neuron pruning is widely used to reduce the computational cost and parameter footprint of large language models, yet it remains unclear whether neurons

model-releasesarxiv-cs-cl
1 May 2026
Model Releases

Fake3DGS: A Benchmark for 3D Manipulation Detection in Neural Rendering

DGX agent

arXiv:2604.27590v1 Announce Type: new Abstract: Recent advances in 3D reconstruction and neural rendering,particularly 3D Gaussian Splatting, make it feasible and simple to edit 3D scenes and re-rende

model-releasesarxiv-cs-cv
1 May 2026
Model Releases

Federated Knowledge Distillation for Multi-Model Architectures Lithography Hotspot Detection

DGX agent

arXiv:2501.04066v2 Announce Type: replace Abstract: As a special type of multimedia data, Lithography Hotspot Detection (LHD) training often requires stronger privacy protection than conventional mult

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

Fidelity, Diversity, and Privacy: A Multi-Dimensional LLM Evaluation for Clinical Data Augmentation

DGX agent

arXiv:2604.27014v1 Announce Type: new Abstract: The scarcity of high-quality annotated medical data, particularly in mental health, poses a significant bottleneck for training robust machine learning

model-releasesarxiv-cs-lg
1 May 2026
Model Releases

FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning

DGX agent

arXiv:2506.02515v4 Announce Type: replace-cross Abstract: Multi-step symbolic reasoning is essential for robust financial analysis; yet, current benchmarks largely overlook this capability. Existing d

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting

DGX agent

arXiv:2604.27974v1 Announce Type: new Abstract: Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluat

model-releasesarxiv-cs-cv
1 May 2026
Model Releases

Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs

DGX agent

arXiv:2506.07180v3 Announce Type: replace-cross Abstract: As video large language models (Video-LLMs) become increasingly integrated into real-world applications that demand grounded multimodal reason

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

FluxMoE: Decoupling Expert Residency for High-Performance MoE Serving

DGX agent

arXiv:2604.02715v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models have become a dominant paradigm for scaling large language models, but their rapidly growing parameter sizes introdu

model-releasesarxiv-cs-lg
1 May 2026
← Previous
1…285286287288289…361
Next →