AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,793 results
Model Releases

Drive-P2D: A Progressive Perception-to-Decision Benchmark for VLMs in Autonomous Driving

DGX agent

arXiv:2601.14702v2 Announce Type: replace Abstract: Autonomous driving requires reliable perception and safe decision-making in complex scenarios. Recent vision-language models (VLMs) demonstrate reas

model-releasesarxiv-cs-ai
27 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

ECHO-2: A Large-Scale Distributed Rollout Framework for Cost-Efficient Reinforcement Learning

DGX agent

arXiv:2602.02192v5 Announce Type: replace Abstract: Reinforcement learning (RL) is a critical stage in post-training large language models (LLMs), involving repeated interaction between rollout genera

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

EpiCurveBench: Evaluating VLMs on Epidemic Curve Digitization

DGX agent

arXiv:2605.27195v1 Announce Type: new Abstract: Chart-to-data extraction with vision-language models (VLMs) is increasingly evaluated on benchmarks that show diminishing headroom (frontier VLMs exceed

model-releasesarxiv-cs-cl
27 May 2026
Research

Error Analysis of Discrete Flow with Generator Matching

DGX agent

arXiv:2509.21906v3 Announce Type: replace-cross Abstract: Discrete flow models offer a powerful framework for learning distributions over discrete state spaces and have demonstrated superior performan

researcharxiv-cs-lg
27 May 2026
Model Releases

InfoSynth: Information-Guided Benchmark Synthesis for LLMs

DGX agent

arXiv:2601.00575v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant advancements in reasoning and code generation, but efficiently creating new benchmarks to

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

LURE: Live-Usage Replay Evaluations for Reducing Evaluation Awareness

DGX agent

arXiv:2605.26438v1 Announce Type: cross Abstract: Large language models can recognize when they are being evaluated (evaluation awareness) and behave differently because of that, which undermines the

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Med-CoReasoner: Reducing Language Disparities in Medical Reasoning via Language-Informed Co-Reasoning

DGX agent

arXiv:2601.08267v3 Announce Type: replace Abstract: While reasoning-enhanced large language models perform strongly on English medical tasks, a persistent multilingual gap remains, with substantially

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

MobileMoE: Scaling On-Device Mixture of Experts

DGX agent

arXiv:2605.27358v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become the de facto architecture for hundred-billion-parameter language models, yet its advantages at sub-billion scales

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

More Expressive Feedforward Layers: Part I. Token-Adaptive Mixing of Activations

DGX agent

arXiv:2605.26647v1 Announce Type: cross Abstract: Feedforward network (FFN) layers account for a large fraction of parameters and nonlinear expressivity in Transformer-based large language models (LLM

model-releasesarxiv-cs-ai
27 May 2026
Research

MSCGC-KAN: Multi-scale Causal Graph Convolution and Kolmogorov-Arnold Feature Mapping for EEG Emotion Recognition

DGX agent

arXiv:2605.26624v1 Announce Type: new Abstract: Electroencephalogram (EEG)-based emotion recognition is an important affective computing task, and recent EEG foundation models provide useful generic r

researcharxiv-cs-cv
27 May 2026
Model Releases

NestedKV: Nested Memory Routing for Long-Context KV Cache Compression

DGX agent

arXiv:2605.26678v1 Announce Type: new Abstract: Long-context language models are limited by the memory footprint of the key-value (KV) cache. Existing training-free KV compression methods usually rank

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

O-MARC: Omni Memory-Augmented Compression Distillation for Efficient Video Understanding

DGX agent

arXiv:2605.26584v1 Announce Type: new Abstract: Omnimodal large language models enable unified audio video understanding, but long joint token sequences make inference costly, and existing benchmarks

model-releasesarxiv-cs-cv
27 May 2026
Local Ai

OmniGF: A Dual-Branch Vision-Language Framework for Unified Gaze Following

DGX agent

arXiv:2605.26399v1 Announce Type: new Abstract: Understanding human gaze behavior is essential for complex scene comprehension and human-computer interaction. Traditional gaze following models are typ

local-aiarxiv-cs-cv
27 May 2026
Tutorials

Optimising Factual Consistency in Summarisation via Preference Learning from Multiple Imperfect Metrics

DGX agent

arXiv:2605.26840v1 Announce Type: new Abstract: Reinforcement learning with evaluation metrics as rewards is widely used to enhance specific capabilities of language models. However, for tasks such as

tutorialsarxiv-cs-cl
27 May 2026
Model Releases

PaTAS: A Framework for Trust Propagation in Neural Networks Using Subjective Logic

DGX agent

arXiv:2511.20586v4 Announce Type: replace Abstract: Trustworthiness has become a key requirement for the deployment of artificial intelligence systems in safety-critical applications. Conventional eva

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Persona2Web: Benchmarking Personalized Web Agents for Contextual Reasoning with User History

DGX agent

arXiv:2602.17003v2 Announce Type: replace-cross Abstract: Large language models have advanced web agents, yet current agents lack personalization capabilities. Since users rarely specify every detail

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Position: AI Safety Requires Effective Controllability

DGX agent

arXiv:2605.27117v1 Announce Type: new Abstract: AI safety is still largely framed as alignment: training models to follow human preferences, safety policies, and normative constraints. That framing ha

model-releasesarxiv-cs-ai
27 May 2026
Agents

Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge Extraction

DGX agent

arXiv:2602.00959v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can be seen as compressed knowledge bases, but it remains unclear what knowledge they truly contain and how far t

agentsarxiv-cs-cl
27 May 2026
Model Releases

SEAL: Self-Evolving Agentic Learning for Conversational Question Answering over Knowledge Graphs

DGX agent

arXiv:2512.04868v2 Announce Type: replace-cross Abstract: Knowledge-based conversational question answering (KBCQA) confronts persistent challenges in resolving coreference, modeling contextual depend

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

SONAR-LLM: Autoregressive Transformer that Thinks in Sentence Embeddings and Speaks in Tokens

DGX agent

arXiv:2508.05305v2 Announce Type: replace Abstract: The recently proposed Large Concept Model (LCM) generates text by predicting a sequence of sentence-level embeddings and training with either mean-s

model-releasesarxiv-cs-cl
27 May 2026
Model Releases

Strategies for Guiding LLMs to Use Software Design Patterns: A Case of Singleton

DGX agent

arXiv:2605.26898v1 Announce Type: cross Abstract: Large Language Models (LLMs) can generate functional source code from natural-language prompts, but often fail to consistently follow higher-level arc

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

Structured Relational Reasoning for Group Activity Assessment

DGX agent

arXiv:2508.07996v2 Announce Type: replace Abstract: Group Activity Detection (GAD) involves recognizing social groups and their collective behaviors in videos. Vision Foundation Models (VFMs), like DI

model-releasesarxiv-cs-cv
27 May 2026
Model Releases

TowerMind: A Tower Defence Game Learning Environment and Benchmark for LLM as Agents

DGX agent

arXiv:2601.05899v2 Announce Type: replace Abstract: Recent breakthroughs in Large Language Models (LLMs) have positioned them as a promising paradigm for agents, with long-term planning and decision-m

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

What Makes Chain-of-Thought Work at Probe Time? Local Co-occurrence Rather Than Global Derivation

DGX agent

arXiv:2605.26795v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting reliably improves language-model accuracy, but which properties of a rationale text drive the improvement is poorly und

model-releasesarxiv-cs-ai
27 May 2026
Model Releases

WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization

DGX agent

arXiv:2605.26660v1 Announce Type: new Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in

model-releasesarxiv-cs-lg
27 May 2026
Model Releases

Acting on the Unseen: Communication-Free Collaborative Filtering for Decentralized Multi-Robot Task Allocation

DGX agent

arXiv:2605.25584v1 Announce Type: cross Abstract: Multi-robot task allocation usually assumes some combination of communication, known task models, or a coordinator. We study the opposite extreme, a r

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

AERIC: Anticipatory Hidden-State Monitoring for Implicit Harmful Dialogue

DGX agent

arXiv:2605.23974v1 Announce Type: new Abstract: Current language models create two safety challenges: risk must be detected early enough to avoid exposing harmful continuation, and the harmfulness its

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

AgentHijack: Benchmarking Computer Use Agent Robustness to Common Environment Corruptions

DGX agent

arXiv:2605.25707v1 Announce Type: new Abstract: Autonomous computer use agents that powered by multimodal large language models (MLLMs) are emerging as capable assistants for completing complex digita

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

An Interactive Paradigm for Deep Research

DGX agent

arXiv:2605.24266v1 Announce Type: cross Abstract: Recent advances in large language models (LLMs) have enabled deep research systems that synthesize comprehensive, report-style answers to open-ended q

model-releasesarxiv-cs-ai
26 May 2026
Research

BackWeak: Backdooring Knowledge Distillation Simply with Weak Triggers and Fine-tuning

DGX agent

arXiv:2511.12046v2 Announce Type: replace-cross Abstract: Knowledge Distillation (KD) is essential for compressing large models, yet relying on pre-trained 'teacher' models downloaded from third-party

researcharxiv-cs-ai
26 May 2026
Model Releases

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

DGX agent

arXiv:2605.24219v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?

DGX agent

arXiv:2605.23913v1 Announce Type: cross Abstract: Cloud-hosted large language models (LLMs) commonly rely on LoRA for domain adaptation, yet domain data are distributed across multiple edge devices an

model-releasesarxiv-cs-cl
26 May 2026
Model Releases

Chain-of-Thought Hijacking

DGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale

DGX agent

arXiv:2605.24305v1 Announce Type: cross Abstract: Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

DGX agent

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Y

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Committed SAE-Feature Traces for Audited-Session Substitution Detection in Hosted LLMs

DGX agent

arXiv:2604.18179v2 Announce Type: replace-cross Abstract: Hosted-LLM providers have a silent-substitution incentive: advertise a stronger model while serving cheaper replies. Probe-after-return scheme

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation

DGX agent

arXiv:2605.25835v1 Announce Type: cross Abstract: This paper examines the specialization of Small Language Models (SLMs) with up to 4 billion parameters for generating artifacts in domain-specific lan

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Continual Speaker Identity Unlearning with Minimal Interference

DGX agent

arXiv:2605.25962v1 Announce Type: cross Abstract: Machine unlearning removes designated concepts or knowledge from pre-trained models. Recent work has extended this paradigm to speaker identity unlear

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Counterfactual Explanations for Hypergraph Neural Networks

DGX agent

arXiv:2602.04360v2 Announce Type: replace-cross Abstract: Hypergraph neural networks (HGNNs) effectively model higher-order interactions in many real-world systems but remain difficult to interpret, l

model-releasesarxiv-cs-ai
26 May 2026
Research

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy

DGX agent

arXiv:2605.25603v1 Announce Type: new Abstract: Chain-of-thought (CoT) reasoning improves the problem-solving ability of large language models (LLMs), but generated reasoning traces may not faithfully

researcharxiv-cs-ai
26 May 2026
Model Releases

E = T*H/(O+B): A Dimensionless Control Parameter for Mixture-of-Experts Ecology

DGX agent

arXiv:2605.06415v2 Announce Type: replace-cross Abstract: We introduce E = T*H/(O+B), a dimensionless control parameter that predicts whether Mixture-of-Experts (MoE) models will develop a healthy exp

model-releasesarxiv-cs-ai
26 May 2026
Local Ai

ExplainReduce: Generating global explanations from many local explanations

DGX agent

arXiv:2502.10311v3 Announce Type: replace-cross Abstract: Most commonly used non-linear machine learning methods are closed-box models, uninterpretable to humans. The field of explainable artificial i

local-aiarxiv-cs-ai
26 May 2026
Model Releases

Fine-Tuning and Serving Gemma 4 31B on Google Cloud TPU: A Technical Comparison with GPU Baselines

DGX agent

arXiv:2605.25645v1 Announce Type: cross Abstract: We present the first end-to-end demonstration of fine-tuning and serving Google's Gemma 4 31B model on TPU hardware, providing an empirical comparison

model-releasesarxiv-cs-ai
26 May 2026
Research

Fine-Tuning Masked Diffusion for Provable Self-Correction

DGX agent

arXiv:2510.01384v4 Announce Type: replace Abstract: A natural desideratum for generative models is self-correction--detecting and revising low-quality tokens at inference. While Masked Diffusion Model

researcharxiv-cs-lg
26 May 2026
Research

Forgettable Federated Linear Learning with Certified Data Unlearning

DGX agent

arXiv:2306.02216v3 Announce Type: replace Abstract: Federated Learning (FL) enables collaborative model training across distributed clients while preserving user privacy. Recently, Federated Unlearnin

researcharxiv-cs-lg
26 May 2026
Model Releases

From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression

DGX agent

arXiv:2605.24316v1 Announce Type: new Abstract: Scaling laws provide compact descriptions of how prediction error varies with compute, model size, and data, but existing theory mainly treats single-sa

model-releasesarxiv-cs-lg
26 May 2026
Model Releases

GlobalDentBench: A Multinational Benchmark for Evaluating LLM Clinical Reasoning in Dentistry with Expert Calibration

DGX agent

arXiv:2605.24636v1 Announce Type: new Abstract: While large language models (LLMs) hold transformative potential for medicine, their reasoning robustness and safety in real-world clinical scenarios re

model-releasesarxiv-cs-ai
26 May 2026
Model Releases

Grammatically-Guided Sparse Attention for Efficient and Interpretable Transformers

DGX agent

arXiv:2605.24518v1 Announce Type: cross Abstract: The quadratic complexity of self-attention in Transformer models remains a significant bottleneck for processing long sequences and deploying large la

model-releasesarxiv-cs-ai
26 May 2026
← Previous
1…461462463464465…1371
Next →