AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
50,764 results
Model Releases

Evaluating Prompt Scope and Demonstration Similarity in Local LLM Machine Translation

DGX agent

arXiv:2607.26286v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as general-purpose translation systems, but their behavior is usually evaluated under a single prompt

model-releasesarxiv-cs-cl
30 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

DGX agent

arXiv:2607.26432v1 Announce Type: new Abstract: Face anti-spoofing (FAS) is increasingly expected to provide not only bona fide/spoof decisions, but also attack semantics and image-grounded evidence f

model-releasesarxiv-cs-cv
30 Jul 2026
Model Releases

From Tokens to Watt-hours: Analytical Energy Estimation for LLM Inference on Modern GPUs

DGX agent

arXiv:2607.26571v1 Announce Type: new Abstract: The operational energy consumption of large language model (LLM) inference is becoming an increasingly important component of the environmental footprin

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

Hearsay: Vision-Language Medical Diagnoses Without an Image

DGX agent

arXiv:2607.26886v1 Announce Type: cross Abstract: When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

InferScale: GPU-Native KV Injection for Personalized LLM Serving

DGX agent

arXiv:2607.27090v1 Announce Type: cross Abstract: Large language models are increasingly deployed with persistent personalized context, such as accumulated memory profiles or long conversation histori

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification

DGX agent

arXiv:2607.26397v1 Announce Type: new Abstract: Enzyme function prediction is a hierarchical, knowledge-intensive form of protein function classification. Existing benchmarks expose an anomaly: genera

model-releasesarxiv-cs-cl
30 Jul 2026
Safety

On-Policy Distillation for LLM Safety: A Routing Approach to Template-Robust Realignment

DGX agent

arXiv:2607.27081v1 Announce Type: cross Abstract: Fine-tuning is the dominant paradigm for specializing large language models (LLMs), yet it exposes a critical vulnerability: malicious data providers

safetyarxiv-cs-cl
30 Jul 2026
Model Releases

ReCo: Reweighting GRPO Against Distributional Concentration

DGX agent

arXiv:2607.26862v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has become a standard reinforcement learning method for post-training language models. Recent work shows that

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

Rethinking Self-Evolution: A Constrained Exploration-Exploitation Process for Mitigating Skill Overfitting

DGX agent

arXiv:2607.26643v1 Announce Type: cross Abstract: Enabling large language model (LLM) agents to accumulate and reuse experience from past interactions remains a central challenge in real-world applica

model-releasesarxiv-cs-lg
30 Jul 2026
Model Releases

SecRespond: Benchmarking AI Agents for Real-World Post-Compromise Incident Response

DGX agent

arXiv:2607.26791v1 Announce Type: cross Abstract: Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs

DGX agent

arXiv:2607.26355v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential

model-releasesarxiv-cs-cl
30 Jul 2026
Model Releases

From Deterministic to Generative Deep Learning for Urban Air Quality Reconstruction from Sparse Observations

DGX agent

arXiv:2607.25687v1 Announce Type: cross Abstract: Full-field reconstruction of air pollution is essential for evaluating pollution exposure and supporting public health decision-making. However, the c

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Generalization from Low- to Moderate-Resolution Spectra with Neural Networks for Stellar Parameter Estimation: A Case Study with DESI

DGX agent

arXiv:2602.15021v2 Announce Type: replace-cross Abstract: Cross-survey generalization is a critical challenge in stellar spectral analysis, particularly in cases such as transferring from low- to mode

model-releasesarxiv-cs-lg
29 Jul 2026
Model Releases

Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

DGX agent

arXiv:2607.24898v1 Announce Type: cross Abstract: State-of-the-art toxicity detectors for text-to-image generation adopt a one-size-fits-all approach: a single universal model applying fixed safety gu

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

Learned, Relied Upon, or Necessary? Separating Checkpoint Dependence from Task-Level Value in Sheaf GNNs

DGX agent

arXiv:2607.25387v1 Announce Type: new Abstract: Learned restriction maps in sheaf graph neural networks are often treated as proof that the model has discovered useful edge geometry. That conclusion d

model-releasesarxiv-cs-lg
29 Jul 2026
Model Releases

MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA

DGX agent

arXiv:2607.24838v1 Announce Type: cross Abstract: In medical multiple-choice question answering (MCQA), Retrieval-Augmented Generation (RAG) can supplement the domain knowledge of language models (LMs

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

MemSFT: Mitigating Alignment Tax with an External Parametric Memory

DGX agent

arXiv:2607.25614v1 Announce Type: cross Abstract: Adapting Large Language Models (LLMs) to specialized domains often incurs an alignment tax, as fine-tuning on domain-specific tasks can cause catastro

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Penelope: Localized Latent Recurrence for Efficient Structured Reasoning

DGX agent

arXiv:2607.25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or b

model-releasesarxiv-cs-ai
29 Jul 2026
Safety

Rashomon Alignment

DGX agent

arXiv:2607.25680v1 Announce Type: cross Abstract: We propose Rashomon Alignment (RA), a new measure to assess functional similarity between two models. Existing functional similarity measures are dist

safetyarxiv-cs-ai
29 Jul 2026
Model Releases

RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation

DGX agent

arXiv:2606.13040v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly explored as visual critics, reward generators, and failure detectors in robotic manipulation. These r

model-releasesarxiv-cs-ro
29 Jul 2026
Model Releases

Shieldstral

DGX agent

arXiv:2607.25857v1 Announce Type: new Abstract: We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7imes its size on text s

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

TabRank: Chain-of-Thought Distillation for Table Re-Rankers

DGX agent

arXiv:2607.25182v1 Announce Type: cross Abstract: The ability to retrieve relevant tables for answering questions is a key task for structured information retrieval. Multi-stage retrieval systems rely

model-releasesarxiv-cs-ai
29 Jul 2026
Model Releases

TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

DGX agent

arXiv:2607.24750v1 Announce Type: new Abstract: Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliabl

model-releasesarxiv-cs-cl
29 Jul 2026
Model Releases

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

DGX agent

arXiv:2607.25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), ho

model-releasesarxiv-cs-ai
29 Jul 2026
Research

A Mechanistic Perspective and Circuit-Guided Difficulty Metric for Unlearning

DGX agent

arXiv:2601.09624v2 Announce Type: replace-cross Abstract: Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably ac

researcharxiv-cs-ai
28 Jul 2026
Model Releases

Adaptive Multi-Scale Forecasting and Gate-Localized Conformal Prediction for Multivariate Nonstationary Time Series

DGX agent

arXiv:2607.23165v1 Announce Type: cross Abstract: We propose ABF-T-GLCP, a model-agnostic framework for forecasting and uncertainty quantification in nonstationary multivariate time series. The centra

model-releasesarxiv-cs-lg
28 Jul 2026
Research

Adversarial Test-Hardening for AI-Written Code: An Instrument Autopsy and a Pre-Registered Causal Estimate of the Critic Loop

DGX agent

arXiv:2607.23002v1 Announce Type: cross Abstract: Large language models increasingly write both code and the tests meant to check it; coverage records what ran, not what was verified. We study an adve

researcharxiv-cs-ai
28 Jul 2026
Model Releases

AlloBench: Measuring Online Tool Allocation Capability in LLM Agents

DGX agent

arXiv:2607.23332v1 Announce Type: new Abstract: Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars

DGX agent

arXiv:2607.24013v1 Announce Type: new Abstract: Production-ready audio-driven avatar generation requires efficient inference without sacrificing fidelity or motion expressiveness. However, existing ac

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

AssumptionMiner: Extracting, Tracing, and Revising Implicit Assumptions in LLM Code Generation

DGX agent

arXiv:2607.22898v1 Announce Type: cross Abstract: Large language models (LLMs) generate code from natural-language prompts, yet real-world prompts rarely provide complete specifications. When prompts

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Between Suppression and Collapse: Evaluating Narrative Unlearning with LENS

DGX agent

arXiv:2607.22657v1 Announce Type: cross Abstract: Large language models (LLMs) can reproduce disinformation-aligned narrative frames as plausible explanations, raising the question of whether existing

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Compressing LLMs with MoP: Mixture of Pruners

DGX agent

arXiv:2602.06127v2 Announce Type: replace Abstract: The high computational demands of Large Language Models (LLMs) motivate methods that reduce parameter count and accelerate inference. In response, m

model-releasesarxiv-cs-lg
28 Jul 2026
Local Ai

Context-Aware Concept Distillation for Trustworthy Flood Prediction

DGX agent

arXiv:2607.23237v1 Announce Type: cross Abstract: Effective flood risk management relies on accurate forecasting, yet the 'black box' nature of stateof-the-art Deep Learning models creates a barrier t

local-aiarxiv-cs-ai
28 Jul 2026
Model Releases

Cost-Aware Recovery-Pathway Identification and Bayesian Optimization for Autonomous Materials Discovery

DGX agent

arXiv:2607.23896v1 Announce Type: new Abstract: Autonomous laboratories automate experimental execution, but a campaign must also decide which recovery pathway merits optimization. We formulate this a

model-releasesarxiv-cs-ai
28 Jul 2026
Safety

Covariance-Boosted Gaussian Processes for Spatiotemporal Irregularities

DGX agent

arXiv:2607.23018v1 Announce Type: cross Abstract: Nonstationary Gaussian process (GP) models are powerful tools for capturing input-dependent variability by adapting to observed data. However, with li

safetyarxiv-cs-lg
28 Jul 2026
Safety

Data Pyramid for Embodied Manipulation

DGX agent

arXiv:2607.24744v1 Announce Type: cross Abstract: Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require d

safetyarxiv-cs-cv
28 Jul 2026
Model Releases

Do LLMs Know Their Vulnerable Scenarios?

DGX agent

arXiv:2607.23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their sa

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

DocHRL: A Hierarchical Reinforcement Learning Framework for Cost-Optimised Document Classification

DGX agent

arXiv:2607.22644v1 Announce Type: new Abstract: Real-world document classification pipelines typically apply the same sequence of models to every incoming document, regardless of its complexity or typ

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

DriveDNA: A Large-Scale Multimodal Naturalistic Driving Dataset and Benchmark for Driving Style Identification

DGX agent

arXiv:2607.23822v1 Announce Type: new Abstract: Driving style captures stable, driver-specific patterns in how a vehicle is driven. In naturalistic data, however, this signal is hard to isolate becaus

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

FilmBench: A Film-Grade Benchmark for Cinematic Video Generation

DGX agent

arXiv:2607.24241v1 Announce Type: cross Abstract: Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw p

model-releasesarxiv-cs-ai
28 Jul 2026
Research

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction

DGX agent

arXiv:2607.23420v1 Announce Type: new Abstract: Large language models show strong promise for information extraction (IE), but existing reflection-based correction methods are often misaligned with st

researcharxiv-cs-cl
28 Jul 2026
Model Releases

Language Shapes Instruction Hierarchy Compliance in Multilingual LLMs

DGX agent

arXiv:2607.23545v1 Announce Type: new Abstract: Instruction hierarchy (IH) requires models to prioritize instructions by source, ensuring that higher-priority instructions override lower-priority ones

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

DGX agent

arXiv:2607.24407v1 Announce Type: new Abstract: Multimodal Large Language Models have made great progress in grounding tasks, yet existing methods still struggle to unify precise localization and comp

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

Multi-Modal Scene Graph with Kolmogorov-Arnold Experts for Audio-Visual Question Answering

DGX agent

arXiv:2511.23304v2 Announce Type: replace Abstract: In this paper, we propose a novel Multi-Modal Scene Graph with Kolmogorov-Arnold Expert Network for Audio-Visual Question Answering (SHRIKE). The ta

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Neuromorphic Object Detection: An In-Depth Study and Future Directions

DGX agent

arXiv:2607.23576v1 Announce Type: new Abstract: Conventional frame-based cameras face significant challenges in detecting objects under high-speed motion blur or in low-light environments. Neuromorphi

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study

DGX agent

arXiv:2410.07809v2 Announce Type: replace Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost. A natural hypothesi

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards

DGX agent

arXiv:2607.23364v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models,

model-releasesarxiv-cs-lg
28 Jul 2026
Model Releases

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries

DGX agent

arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions

model-releasesarxiv-cs-ai
28 Jul 2026
← Previous
1…305306307308309…1058
Next →