AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,619
  • Agents7,270
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,100
  • Local Ai4,731
  • Model Releases22,595
  • Research19,194
  • Safety12,820
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
84,619Total entries
1Added by human
84,618Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
17,288 results
Model Releases

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

DGX agent

arXiv:2605.25510v1 Announce Type: new Abstract: Children increasingly have access to Large Language Models (LLMs), which may expose them to responses that are developmentally inappropriate or require

model-releasesarxiv-cs-cl
26 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

The LSCD Benchmark: a Testbed for Diachronic Word Meaning Tasks

DGX agent

arXiv:2404.00176v3 Announce Type: replace Abstract: Lexical Semantic Change Detection (LSCD) is a complex, lemma-level task, which is usually operationalized based on two subsequently applied usage-le

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

The Model Is Not the Product: A Dual-Pillar Architecture for Local-First Psychological Coaching

DGX agent

arXiv:2605.24411v1 Announce Type: new Abstract: Existing language model applications struggle to meet the demand for emotionally oriented support, primarily due to their inability to maintain deep, pe

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

The Perception-Physics Paradox: Probing Scientific Alignment with TC-Bench

DGX agent

arXiv:2605.24782v1 Announce Type: new Abstract: While Vision Foundation Models (VFMs) excel at predictive tasks on satellite imagery, their performance can arise from visual correlations rather than u

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

The Time is Here for Just-in-Time Systems: Challenges and Opportunities

DGX agent

arXiv:2605.24096v1 Announce Type: cross Abstract: Core systems like key-value stores have historically taken years to build, and are designed to be general so as to amortize cost across deployments, p

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning

DGX agent

arXiv:2605.25850v1 Announce Type: cross Abstract: This paper investigates large language model (LLM) abstention learning, specifically using ternary reward, which incentivize truthfulness in large lan

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

TimeSpot: Benchmarking Geo-Temporal Understanding in Vision-Language Models in Real-World Settings

DGX agent

arXiv:2603.06687v2 Announce Type: replace-cross Abstract: Geo-temporal understanding, the ability to infer location, time, and contextual properties from visual input alone, underpins applications suc

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Tiny Brains, Giant Impact: Uncovering the Keystone Neurons of LLM with Just a Few Prompts

DGX agent

arXiv:2605.24846v1 Announce Type: cross Abstract: Large language models (LLMs) display strong comprehensive abilities, yet the internal mechanisms that support these behaviors remain insufficiently un

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ToolRegistry: A Protocol-Agnostic Tool Management Library for Function-Calling LLMs

DGX agent

arXiv:2507.10593v3 Announce Type: replace-cross Abstract: Every LLM tool call is structurally an RPC -- a function name, JSON arguments, and a serialized result -- yet each protocol (native Python, MC

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Topology-Driven Transferability Estimation of Medical Foundation Models for Segmentation

DGX agent

arXiv:2602.23916v2 Announce Type: replace-cross Abstract: The advent of large-scale self-supervised learning (SSL) has produced a vast zoo of medical foundation models. However, selecting optimal medi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Toward a Benchmark for Controllable Simulation of Imperfect Students with Large Language Models

DGX agent

arXiv:2605.25601v1 Announce Type: cross Abstract: Teacher education requires deliberate practice with learners who exhibit identifiable strengths, weaknesses, and partial mastery. Large language model

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Towards Large Model Feature Coding

DGX agent

arXiv:2605.24025v1 Announce Type: cross Abstract: Large models have delivered remarkable performance across a wide range of perception and generation tasks, yet practical deployment is increasingly co

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

TRACE: A taxonomy-grounded synthetic dataset for teaching-program generation and session interpretation in Applied Behavior Analysis

DGX agent

arXiv:2605.25038v1 Announce Type: new Abstract: Applied Behavior Analysis (ABA) is a clinical discipline whose documentation, teaching programs and multi-session behavioral logs, is formulaic and high

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs

DGX agent

arXiv:2605.24079v1 Announce Type: cross Abstract: Data contamination is a known threat to the reliability of model evaluation. However, it remains underexplored in code large language models (LLMs), w

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Trade-off Functions for DP-SGD with Subsampling based on Random Shuffling: Tight Upper and Lower Bounds

DGX agent

arXiv:2605.06259v2 Announce Type: replace Abstract: We derive a tight analysis of the trade-off function for Differentially Private Stochastic Gradient Descent (DP-SGD) with subsampling based on rando

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

TriVAL: A Tri-Validation Framework for Faithful Automatic Optimization Modeling

DGX agent

arXiv:2605.23966v1 Announce Type: cross Abstract: Optimization modeling serves as the pivotal bridge between natural-language problem descriptions and optimization solvers, and remains a cornerstone f

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Trust but Verify: Prover-Verifier Deliberation for Selective LLM Prediction

DGX agent

arXiv:2605.25133v1 Announce Type: new Abstract: Reliably knowing when a language model is correct is almost as important as being correct. We introduce prover-verifier deliberation (PVD), an inference

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Truthful Online Preference Aggregation for LLM Fine-Tuning in Mobile Crowdsourcing

DGX agent

arXiv:2605.24052v1 Announce Type: cross Abstract: To better serve users' demands in mobile applications (e.g., navigation), mobile crowdsourcing platforms can iteratively align large language model (L

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering

DGX agent

arXiv:2605.24703v1 Announce Type: cross Abstract: Large language models (LLMs) and time-series language models (TSLMs) are increasingly applied to time-series question answering (TSQA). Unlike text-on

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

TTPrint: Evidence-Grounded TTP Extraction via Diverge-then-Converge Verification

DGX agent

arXiv:2605.25836v1 Announce Type: cross Abstract: Extracting MITRE ATT&CK techniques from cyber threat intelligence (CTI) reports is an open-set, multi-label problem requiring both high recall (not mi

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

TypedCSIP: Typed Counterfactual Pretraining for Chinese Legislative Conflict Classification

DGX agent

arXiv:2605.25474v1 Announce Type: new Abstract: TypedCSIP is a typed counterfactual pretraining method for the conflict-classification task of the LCR-CN benchmark (Zhao et al., 2026): given a (superi

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Uncertainty Decomposition via Cyclical SG-MCMC and Soft-label Learning for Subjective NLP

DGX agent

arXiv:2605.24773v1 Announce Type: new Abstract: Annotator disagreement in emotion classification reflects ambiguity intrinsic to emotion concepts and is essential for predictor-quality assessment in s

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Understanding and Mitigating Premature Confidence for Better LLM Reasoning

DGX agent

arXiv:2605.24396v1 Announce Type: new Abstract: Long chains of thought (CoT) from current language models frequently contain logical gaps and unjustified leaps, limiting the gains from additional test

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Understanding Conversational Patterns in Multi-agent Programming: A Case Study on Fibonacci Game Development

DGX agent

arXiv:2605.24138v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly applied to software engineering (SE), yet their potential for autonomous, role-oriented collaboration re

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Uni-DPO: A Unified Paradigm for Dynamic Preference Optimization of LLMs

DGX agent

arXiv:2506.10054v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO) has emerged as a cornerstone of reinforcement learning from human feedback (RLHF) due to its simplicity a

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

URS: A Unified Neural Routing Solver for Cross-Problem Zero-Shot Generalization

DGX agent

arXiv:2509.23413v2 Announce Type: replace Abstract: Multi-task neural routing solvers have emerged as a promising paradigm for their ability to solve multiple vehicle routing problems (VRPs) using a s

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

UtilityMax Prompting: A Formal Framework for Multi-Objective Large Language Model Tasks

DGX agent

arXiv:2603.11583v4 Announce Type: replace-cross Abstract: The success of a Large Language Model (LLM) task depends heavily on its prompt. Most use-cases specify prompts using natural language, which i

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

UWM-JEPA: Predictive World Models That Imagine in Belief Space

DGX agent

arXiv:2605.25313v1 Announce Type: cross Abstract: World models for partially observed environments must imagine multiple compatible hidden futures and steer between them under counterfactual actions.

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

V3H: View Variation and View Heredity for Incomplete Multi-view Clustering

DGX agent

arXiv:2011.11194v4 Announce Type: replace Abstract: Real data often appear in the form of multiple incomplete views. Incomplete multi-view clustering is an effective method to integrate these incomple

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

VaaWIT: Visual-Aware Adaptation of Large Language Models for Multilingual Web Image Translation

DGX agent

arXiv:2605.24675v1 Announce Type: cross Abstract: Translating text embedded in Web images is crucial for improving content accessibility and cross-lingual information retrieval, particularly within so

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

vAttention: Verified Sparse Attention

DGX agent

arXiv:2510.05688v2 Announce Type: replace-cross Abstract: State-of-the-art sparse attention methods for reducing decoding latency fall into two main categories: approximate top-k (and its extension, t

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

VeriTrace: Evolving Mental Models for Deep Research Agents

DGX agent

arXiv:2605.26081v1 Announce Type: new Abstract: Deep research agents face vast, interdependent, and pervasively uncertain information. Existing systems explore what evolving intermediate representatio

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ViroBench: Benchmarking Nucleotide Foundation Models on Viral Genomics Tasks

DGX agent

arXiv:2605.25388v1 Announce Type: new Abstract: Nucleotide sequences constitute the fundamental genetic basis of biological systems, rendering viral genomic analysis critical for biomedical advancemen

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models

DGX agent

arXiv:2605.25820v1 Announce Type: new Abstract: Diffusion-based multimodal large language models (dMLLMs) decode by iteratively predicting tokens at multiple masked positions in parallel. This turns e

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes

DGX agent

arXiv:2509.25339v3 Announce Type: replace-cross Abstract: Is basic visual understanding really solved in state-of-the-art VLMs? We present VisualOverload, a slightly different visual question answerin

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

What Makes a Medical Checker Trainable? Diagnosing Signal Collapse and Reward Hacking in Checker-Guided RAG for Biomedical QA

DGX agent

arXiv:2605.25988v1 Announce Type: new Abstract: Medical RAG needs evidence-grounded claims, so plugging a claim-level NLI checker into retrieval-augmented RL is intuitive. extbf{We find that the check

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

When Can We Trust Early Warnings? Leakage-Excluded Early Outcome Prediction from LMS Interaction Logs

DGX agent

arXiv:2605.25794v1 Announce Type: new Abstract: Early-warning models built from Learning Management System (LMS) logs aim to predict end-of-course outcomes early enough to enable timely learner suppor

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure

DGX agent

arXiv:2605.23932v1 Announce Type: new Abstract: Despite strong medical benchmark accuracy, LLMs can exhibit severe multi-turn sycophancy in clinical dialogue, abandoning initial correct diagnosis unde

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Do LLM Agents Treat Surface Noise Differently from Semantic Noise? A 68-Cell Measurement Study with a Held-Out Trace-Level Validation

DGX agent

arXiv:2605.25981v1 Announce Type: new Abstract: We document an empirical phenomenon in chain-of-thought and ReAct agents driven by ten large language models from seven architecture families: meaning-b

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

When Mean CE Fails: Median CE Can Better Track Language Model Quality

DGX agent

arXiv:2605.24667v1 Announce Type: new Abstract: Mean cross-entropy is the standard validation metric for language models, but it can fail to track model quality during training. We examine this in two

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Reasoning Hurts: Source-Aware Evaluation of Frontier LLMs for Clinical SOAP Note Generation

DGX agent

arXiv:2605.24902v1 Announce Type: cross Abstract: Reasoning-enabled LLMs perform strongly on medical reasoning benchmarks, but it remains unclear whether these gains transfer to structured clinical do

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When Search Becomes Memory: Turning Robot Design Trials into Transferable Skills

DGX agent

arXiv:2605.25832v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as proposal generators for evolutionary robot design, yet most loops remain memoryless: simulator r

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

When the Manual Lies: A Realistic Benchmark to Evaluate MCP Poisoning Attacks for LLM Agents

DGX agent

arXiv:2605.24069v1 Announce Type: cross Abstract: The rise of tool-using Large Language Model (LLM) agents, standardized by protocols like the Model Context Protocol (MCP), has unlocked unprecedented

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

WhenLoss: Diagnosing Write and Retrieval Bottlenecks in Long-Context Memory Systems

DGX agent

arXiv:2605.24579v1 Announce Type: new Abstract: Long-context memory systems often fail under fixed budgets, but end-to-end evaluation does not reveal whether evidence was discarded during compression

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring

DGX agent

arXiv:2605.24737v1 Announce Type: cross Abstract: Current approaches to AI compliance treat conformity as a binary, audit-time verdict rather than a continuous, measurable property of production syste

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

WhoSaidIt: Human-LLM Collaborative Annotation for Text-Based Multilingual Speaker-Attribute Classification

DGX agent

arXiv:2605.26070v1 Announce Type: new Abstract: Annotating speaker attributes from text is inherently ambiguous, particularly in multilingual settings where demographic and social cues are implicit an

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

DGX agent

arXiv:2605.25256v1 Announce Type: new Abstract: Aligning AI systems with organizational decision-making is typically framed as a single-target problem: make the model behave like the organization. We

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

WideDepth: Millimeter-Accurate Benchmark for Fisheye Depth Estimation

DGX agent

arXiv:2605.24074v1 Announce Type: cross Abstract: Fisheye cameras are increasingly adopted in robotics for near-field manipulation, navigation, and immersive perception, yet indoor depth benchmarks wi

model-releasesarxiv-cs-ro
26 May 2026
← Previous
1…203204205206207…361
Next →