AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,630
  • Agents7,271
  • Applications5,200
  • Concepts5
  • Hardware1,757
  • Industry6,101
  • Local Ai4,731
  • Model Releases22,603
  • Research19,194
  • Safety12,821
  • Syntheses17
  • Tools1,668
  • Tutorials3,262

Source
HumanDGX agent
84,630Total entries
1Added by human
84,629Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,603 results
26 May 2026

Beyond Literal Translation: Evaluating Cultural Effectiveness in Social Media UGC

Model ReleasesDGX agent

arXiv:2605.25626v1 Announce Type: new Abstract: Social media platforms enable large-scale cross-lingual communication, but translating user-generated content (UGC) remains challenging due to its infor

Beyond Query Memorization: Large Language Model Routing with Query Decomposition and Historical Matching

Model ReleasesDGX agent

arXiv:2605.25558v1 Announce Type: new Abstract: Optimizing the trade-off among predictive performance and computational cost is a central focus in the deployment of Large Language Models (LLMs). Curre

Beyond Summaries: Structure-Aware Labeling of Code Changes with Large Language Models

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2605.26100v1 Announce Type: cross Abstract: Code review is a critical practice in software engineering, yet the growing scale and frequency of code patches in modern projects, together with the

BODHI: Precise OS Kernel Specification Inference

Model ReleasesDGX agent

arXiv:2605.23931v1 Announce Type: new Abstract: The formal verification of operating system kernels requires precise specifications that capture the intended behavior of system calls. Writing these sp

Bridging On-Device and Cloud LLMs for Collaborative Reasoning: A Unified Methodology for Local Routing and Post-Training

Model ReleasesDGX agent

arXiv:2509.24050v4 Announce Type: replace Abstract: Device-cloud collaboration holds promise for deploying large language models (LLMs), leveraging lightweight on-device models for efficiency while re

Building an Adversarial Malware Dataset by Family and Type: Generation, Evasion, and Poisoning Evaluation

Model ReleasesDGX agent

arXiv:2605.25937v1 Announce Type: cross Abstract: We present a dataset of adversarial malware samples derived from the public RawMal-TF collection of real-world malware binaries. Using a suite of adve

Can LLMs Time Travel? Enhancing Temporal Consistency in Legal Agentic Search through Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.25920v1 Announce Type: cross Abstract: While large language models (LLMs) augmented with agentic search capabilities show promise for legal reasoning, they overlook a fundamental constraint

Can LoRA Fusion Support Cross-Domain Tasks in Cloud-Edge Collaboration?

Model ReleasesDGX agent

arXiv:2605.23913v1 Announce Type: cross Abstract: Cloud-hosted large language models (LLMs) commonly rely on LoRA for domain adaptation, yet domain data are distributed across multiple edge devices an

Cascade-KDE: Robust Time-Series Restoration under Out-of-Distribution Impulse Corruptions

Model ReleasesDGX agent

arXiv:2605.24055v1 Announce Type: cross Abstract: Real-world time-series data in industrial sensing, healthcare, and energy systems is often corrupted by a mixture of Gaussian noise and occasional lar

Causal Tongue-Tie: LLMs Can Encode Causal Direction, But Their Yes/No Outputs Fail to Express

Model ReleasesDGX agent

arXiv:2605.25891v1 Announce Type: cross Abstract: We find a mismatch between what large language models encode about a causal question and what they answer. On anti-commonsense CLadder items, a fixed

CausaLab: A Scalable Environment for Interactive Causal Discovery Toward AI Scientists

Model ReleasesDGX agent

arXiv:2605.26029v1 Announce Type: new Abstract: We introduce CausaLab, a scalable environment for evaluating interactive causal discovery by LLM agents. Unlike prior evaluations, CausaLab evaluates bo

Chain-of-Thought Hijacking

Model ReleasesDGX agent

arXiv:2510.26418v4 Announce Type: replace Abstract: Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reas

ChainLearn: A Blockchain-Based Capacity-Aware Framework for Federated Ensemble Learning

Model ReleasesDGX agent

arXiv:2605.24418v1 Announce Type: new Abstract: Federated learning is used in medical imaging where privacy prohibits centralizing data. Standard federated algorithms assume homogeneous hardware, iden

ChaosBench-Logic v2: Evaluating LLM Logical Reasoning over Dynamical Systems at Scale

Model ReleasesDGX agent

arXiv:2605.24305v1 Announce Type: cross Abstract: Standard accuracy on binary reasoning benchmarks hides critical failure modes: prior collapse, inconsistency under paraphrase, and inability to reason

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

Model ReleasesDGX agent

arXiv:2510.02361v2 Announce Type: replace-cross Abstract: Transformer-based large models excel in natural language processing and computer vision, but face severe computational inefficiencies due to t

CITYREP: A Unified Benchmark for Urban Representations Across Cities, Tasks, and Modalities

Model ReleasesDGX agent

arXiv:2605.26036v1 Announce Type: new Abstract: Urban representation learning encodes complex urban environments into general-purpose embeddings for diverse downstream tasks and emerging urban foundat

Clarification Is Not Enough: Post-Clarification Answering Remains the Bottleneck in Multi-Turn QA

Model ReleasesDGX agent

arXiv:2605.25204v1 Announce Type: new Abstract: Pluralistic alignment requires systems to adapt to diverse user values, communication styles, and contextual assumptions. We believe that a foundational

Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World

Model ReleasesDGX agent

arXiv:2605.26086v1 Announce Type: new Abstract: Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Y

CMAP: Cross-Modal Adaptive Prompting for Multi-Domain Task-Incremental Learning

Model ReleasesDGX agent

arXiv:2605.25708v1 Announce Type: cross Abstract: Multi-domain task-incremental learning requires a model to sequentially acquire knowledge across visually diverse domains without forgetting prior tas

Code2UML: Agentic LLMs with context engineering for scalable software visualization

Model ReleasesDGX agent

arXiv:2605.24453v1 Announce Type: cross Abstract: Large Language Model (LLM)-based code analysis tools are adopted to automate software documentation tasks. However, the scalability of these approache

CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation

Model ReleasesDGX agent

arXiv:2605.25378v1 Announce Type: cross Abstract: Customized image editing aims to equip pre-trained diffusion models with specific visual effects using limited paired data, typically via Low-Rank Ada

Committed SAE-Feature Traces for Audited-Session Substitution Detection in Hosted LLMs

Model ReleasesDGX agent

arXiv:2604.18179v2 Announce Type: replace-cross Abstract: Hosted-LLM providers have a silent-substitution incentive: advertise a stronger model while serving cheaper replies. Probe-after-return scheme

Complement Submodular Information Measures for Balanced and Robust Data Selection

Model ReleasesDGX agent

arXiv:2605.24779v1 Announce Type: cross Abstract: Submodular optimization has become a fundamental paradigm for data selection, retrieval, summarization, and representation learning due to its ability

Concept Unlearning via Cross-Attention Activation Projection for Diffusion Models

Model ReleasesDGX agent

arXiv:2605.25765v1 Announce Type: cross Abstract: Concept unlearning aims to erase a target concept from a pretrained text-to-image diffusion model without retraining. Closed-form methods are attracti

Conformalised imprecise inference for robust extrapolation under limited data

Model ReleasesDGX agent

arXiv:2605.25882v1 Announce Type: new Abstract: Recent advances in uncertainty quantification increasingly emphasise the distinction between aleatory and epistemic uncertainty in machine learning, mot

Context-Instrumental Data Distillation for Kubernetes Manifest Generation: Method and Experimental Evaluation

Model ReleasesDGX agent

arXiv:2605.25835v1 Announce Type: cross Abstract: This paper examines the specialization of Small Language Models (SLMs) with up to 4 billion parameters for generating artifacts in domain-specific lan

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

Model ReleasesDGX agent

arXiv:2605.24279v1 Announce Type: new Abstract: A frontier language model's acknowledged 'helpful programming assistant' persona does not survive long agentic-coding sessions in the deployment regime

Continual Speaker Identity Unlearning with Minimal Interference

Model ReleasesDGX agent

arXiv:2605.25962v1 Announce Type: cross Abstract: Machine unlearning removes designated concepts or knowledge from pre-trained models. Recent work has extended this paradigm to speaker identity unlear

Convex-Neural RRT*: Fast and Reliable Learning-Guided Sampling for High-Quality Robot Path Planning

Model ReleasesDGX agent

arXiv:2605.25006v1 Announce Type: cross Abstract: Sampling-based algorithms for robot path planning offer probabilistic completeness and strong empirical convergence properties across environments wit

Counterfactual Explanations for Hypergraph Neural Networks

Model ReleasesDGX agent

arXiv:2602.04360v2 Announce Type: replace-cross Abstract: Hypergraph neural networks (HGNNs) effectively model higher-order interactions in many real-world systems but remain difficult to interpret, l

Courtroom Analogy: New Perspective on Uncertainty-Aware Classification

Model ReleasesDGX agent

arXiv:2605.25616v1 Announce Type: new Abstract: Single-pass uncertainty quantification (UQ) methods for classification represent uncertainty by predicting a tractable distribution over the class proba

Critical Organization of Deep Neural Networks, and p-Adic Statistical Field Theories

Model ReleasesDGX agent

arXiv:2601.19070v2 Announce Type: replace Abstract: We rigorously study the thermodynamic limit of deep neural networks (DNNS) and recurrent neural networks (RNNs), assuming that the activation functi

Cross-Domain Energy-Guided Diffusion Generation for Off-Dynamics Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.24810v1 Announce Type: cross Abstract: Off-dynamics offline reinforcement learning seeks to learn a target-domain policy from a large source dataset and a limited target dataset under misma

Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection

Model ReleasesDGX agent

arXiv:2605.24965v1 Announce Type: cross Abstract: The rapid evolution of generative models has enabled the creation of hyper-realistic facial deepfakes, exposing a critical vulnerability in modern dig

CSP-Atlas: Concept-Specific Neural Circuits in a Sparse Python Transformer

Model ReleasesDGX agent

arXiv:2605.24603v1 Announce Type: new Abstract: A sparse 8-layer code transformer develops dedicated neural circuitry for every Python construct tested, and that circuitry is organised by a clean comp

CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents

Model ReleasesDGX agent

arXiv:2605.25624v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has driven breakthroughs in domains such as math, tool-use, and software engineering, yet its exte

CurveRL: Principled Distribution-Aware Context Reweighting for LLM Reasoning

Model ReleasesDGX agent

arXiv:2605.24331v1 Announce Type: new Abstract: Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the

CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering

Model ReleasesDGX agent

arXiv:2605.24765v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly applied to cybersecurity question answering (QA) for critical tasks such as incident response and vulner

D^2-Monitor: Dynamic Safety Monitoring for Diffusion LLMs via Hesitation-Aware Routing

Model ReleasesDGX agent

arXiv:2605.25893v1 Announce Type: new Abstract: Despite the emergence of diffusion large language models (D-LLMs) as an alternative to autoregressive large language models (AR-LLMs), safety monitoring

DarkForest: Less Talk, Higher Accuracy for Multi-Agent LLMs

Model ReleasesDGX agent

arXiv:2605.25188v1 Announce Type: new Abstract: Multi-agent LLM systems improve reasoning by combining outputs from multiple agents, but interaction-heavy methods can introduce error propagation and h

Data-Specific Hyper-Parameter Design: A Paradigm Shift in Reservoir Computing

Model ReleasesDGX agent

arXiv:2605.25221v1 Announce Type: cross Abstract: Reservoir computing typically relies on large, randomly generated reservoirs, enabling simple, often linear readouts. Over the past two decades, most

Decision-Making with Lightweight Confidence-Aware Language Model for Autonomous Driving

Model ReleasesDGX agent

arXiv:2605.25393v1 Announce Type: new Abstract: Large Language Models (LLMs) and Multimodal LLMs (MLLMs) have demonstrated immense potential in autonomous driving (AD) by offering human-like reasoning

Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval

Model ReleasesDGX agent

arXiv:2605.24454v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong performance in the legal domain, demonstrating notable potential in Legal Question Answering (LQA). Howev

Deployment-complete benchmarking

Model ReleasesDGX agent

arXiv:2605.25997v1 Announce Type: new Abstract: Benchmarks increasingly guide deployment, procurement and scientific screening, yet a score supports only the response it records, not necessarily the d

Directional Alignment Mitigates Reward Hacking in Reinforcement Learning for Language Models

Model ReleasesDGX agent

arXiv:2605.25189v1 Announce Type: cross Abstract: Reward hacking arises when a model improves a proxy reward by exploiting shortcuts rather than solving the intended task. We study this failure mode t

DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

Model ReleasesDGX agent

arXiv:2605.26087v1 Announce Type: cross Abstract: Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of establis

Distilling Game Code World Model Generation into Lightweight Large Language Models

Model ReleasesDGX agent

arXiv:2605.24375v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown great ability in generating executable code from natural language, opening the possibility of automatically cons

Distributionally Robust Transfer Learning with Structurally Missing Covariates, with Application to Cross-National Cardiac Arrest Prediction

Model ReleasesDGX agent

arXiv:2605.24212v1 Announce Type: cross Abstract: Deploying clinical prediction models across healthcare systems often fails when key training covariates are unavailable at deployment and labeled outc

DIVER-1: Scaling Intracranial EEG Foundation Models for Transferable Representations

Model ReleasesDGX agent

arXiv:2512.19097v3 Announce Type: replace-cross Abstract: Intracranial EEG (iEEG) provides direct, millisecond-scale recordings of human neural activity, but reusable representation learning is diffic

Double Triangle Annotation: A Scalable Human-in-the-Loop Framework for High-Precision Historical Document Annotation

Model ReleasesDGX agent

arXiv:2605.25781v1 Announce Type: new Abstract: Evaluating structured-information extraction from historical documents at scale requires high-precision ground-truth annotations, yet traditional manual

DRInQ: Evaluating Conversational Implicature with Controlled Context Variation

Model ReleasesDGX agent

arXiv:2605.24267v1 Announce Type: new Abstract: Human conversation relies heavily on conversational implicature, in which speakers convey meanings that are suggested rather than explicitly stated. Alt

DropoutTS: Sample-Adaptive Dropout for Robust Time Series Forecasting

Model ReleasesDGX agent

arXiv:2601.21726v2 Announce Type: replace Abstract: Deep time series models are vulnerable to noisy data ubiquitous in real-world applications. Existing robustness strategies either prune data or rely

DRScaffold: Boosting Dense-Scene Reasoning in Lightweight Vision Language Models

Model ReleasesDGX agent

arXiv:2605.26038v1 Announce Type: cross Abstract: Lightweight vision-language models perform competitively on standard benchmarks yet fail systematically in dense-scene reasoning, where multiple objec

Dynamics Reveals Structure: Challenging the Linear Propagation Assumption

Model ReleasesDGX agent

arXiv:2601.21601v2 Announce Type: replace-cross Abstract: Neural networks adapt through first-order parameter updates, yet it remains unclear whether such updates preserve logical coherence. We invest

DynaPURLS: Dynamic Refinement of Part-Aware Representations for Skeleton-Based Zero-Shot Action Recognition

Model ReleasesDGX agent

arXiv:2512.11941v2 Announce Type: replace-cross Abstract: Zero-shot skeleton-based action recognition (ZS-SAR) is fundamentally constrained by prevailing approaches that rely on aligning skeleton feat

E = T*H/(O+B): A Dimensionless Control Parameter for Mixture-of-Experts Ecology

Model ReleasesDGX agent

arXiv:2605.06415v2 Announce Type: replace-cross Abstract: We introduce E = T*H/(O+B), a dimensionless control parameter that predicts whether Mixture-of-Experts (MoE) models will develop a healthy exp

EchoDistill:Alignment Noisy-to-Clean Self-Distillation for Robust Audio LLMs

Model ReleasesDGX agent

arXiv:2605.23954v1 Announce Type: cross Abstract: Audio Large Language Models (ALLMs) are highly vulnerable to real-world noise, which often induces severe semantic drift and hallucinations. Existing

EchoPilot: Training-Free Ultrasound Video Segmentation via Scale-Space Semantic Prompting and Reliability-Gated Memory

Model ReleasesDGX agent

arXiv:2605.25944v1 Announce Type: cross Abstract: Ultrasound video segmentation is clinically valuable yet difficult due to speckle noise, weak boundaries, and rapid anatomical deformation. Recent pro

Efficient Benchmarking Is Just Feature Selection and Multiple Regression

Model ReleasesDGX agent

arXiv:2605.25773v1 Announce Type: cross Abstract: Efficient benchmarking techniques aim to lower the computational cost of evaluating LLMs by predicting full benchmark scores using only a subset of a

Efficient DP-SGD for LLMs with Randomized Clipping

Model ReleasesDGX agent

arXiv:2605.24879v1 Announce Type: new Abstract: Large language models (LLMs) are trained on vast datasets that may contain sensitive information. Differential privacy (DP), the de facto standard for f

← Previous
1…209210211212213…377
Next →