AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
90,914Total entries
1Added by human
90,913Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
53,690 results
Model Releases

Aligned but Fragile: Enhancing LLM Safety Robustness via Zeroth-Order Optimization

DGX agent

arXiv:2605.29396v1 Announce Type: new Abstract: Safety alignment for large language models (LLMs) aims to reduce harmful or unsafe behavior while preserving general utility. However, recent findings r

model-releasesarxiv-cs-ai
29 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Beyond Recall: Behavioral Specification as an Interpretive Layer for AI Personalization

DGX agent

arXiv:2605.28969v1 Announce Type: cross Abstract: If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how f

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Brain-IT-VQA: From Brain Signals to Answers

DGX agent

arXiv:2605.29588v1 Announce Type: cross Abstract: Decoding visual content from fMRI signals recorded while a person views images, and specifically answering questions about the seen images, is a long-

model-releasesarxiv-cs-ai
29 May 2026
Safety

Causal Interventions on Continuous Variables: A Case Study on Verb Bias in Steering Vectors for In-Context Learning

DGX agent

arXiv:2605.29971v1 Announce Type: new Abstract: Causal interventions in language model representations have largely targeted discrete features, like grammatical number. However, language models must a

safetyarxiv-cs-cl
29 May 2026
Safety

CB-SLICE: Concept-Based Interpretable Error Slice Discovery

DGX agent

arXiv:2605.29836v1 Announce Type: cross Abstract: Despite strong average-case performance, deep learning models often exhibit systematic errors on specific population groups, known as error slices. Id

safetyarxiv-cs-ai
29 May 2026
Model Releases

Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering

DGX agent

arXiv:2605.29742v1 Announce Type: new Abstract: Deploying Large Language Models (LLMs) for regulatory compliance demands rigorous traceability via comprehensive citations across multi-tiered authority

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization

DGX agent

arXiv:2510.14150v5 Announce Type: replace Abstract: We introduce CodeEvolve, an open-source framework that couples large language models with island-based evolutionary search for end-to-end algorithmi

model-releasesarxiv-cs-ai
29 May 2026
Research

Collaborative Threshold Watermarking

DGX agent

arXiv:2602.10765v2 Announce Type: replace Abstract: In federated learning (FL), K clients jointly train a model without sharing raw data. Because each participant invests data and compute, clients nee

researcharxiv-cs-lg
29 May 2026
Model Releases

Combating Data Laundering in LLM Training

DGX agent

arXiv:2604.01904v2 Announce Type: replace-cross Abstract: Data rights owners can detect unauthorized data use in large language model (LLM) training by querying with proprietary samples. Often, superi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective

DGX agent

arXiv:2502.03805v2 Announce Type: replace Abstract: Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Diffusion-based learning framework for Constrained Nonconvex Optimization with Weighted Bootstrapped Refinement

DGX agent

arXiv:2502.10330v4 Announce Type: replace Abstract: Recent advances in diffusion models show promising potential to accelerate nonconvex problem solving by leveraging their multimodality. However, mos

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Dynamics of Stochastic Momentum with Sparse Updates in High Dimensions

DGX agent

arXiv:2605.28961v1 Announce Type: cross Abstract: Existing theory of momentum assumes that gradients arrive at every parameter at a roughly constant rate, an assumption violated in practice by heavy-t

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

DySem: Uncovering Dynamic Semantic Components via Multilingual Consensus for Calculating Semantic Textual Similarity

DGX agent

arXiv:2605.29751v1 Announce Type: new Abstract: Calculating semantic textual similarity is a foundational task in natural language processing. Current large language models (LLMs) based methods typica

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Empathic Prompting: Non-Verbal Context Integration for Multimodal LLM Conversations

DGX agent

arXiv:2510.20743v2 Announce Type: replace-cross Abstract: We present Empathic Prompting, a novel framework for multimodal human-AI interaction that enriches Large Language Model (LLM) conversations wi

model-releasesarxiv-cs-ai
29 May 2026
Safety

EvoRubric: Self-Evolving Rubric-Driven RL for Open-Ended Generation

DGX agent

arXiv:2605.29847v1 Announce Type: new Abstract: Reinforcement Learning (RL) has significantly advanced Large Language Models (LLMs) in verifiable domains, but aligning models for open-ended generation

safetyarxiv-cs-cl
29 May 2026
Model Releases

ExCAM: Explainable Cultural Awareness Metrics

DGX agent

arXiv:2605.29897v1 Announce Type: new Abstract: Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications

model-releasesarxiv-cs-cl
29 May 2026
Model Releases

Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning

DGX agent

arXiv:2602.01058v2 Announce Type: replace-cross Abstract: Post-training of reasoning LLMs is a holistic process that typically consists of an offline SFT stage followed by an online reinforcement lear

model-releasesarxiv-cs-ai
29 May 2026
Safety

GRUFF: LLM Pronoun Fidelity, Reasoning, and Biases in German

DGX agent

arXiv:2605.30214v1 Announce Type: new Abstract: Third-person singular pronouns have long been used to study stereotypical biases in language models and to test their abilities to reason about referenc

safetyarxiv-cs-cl
29 May 2026
Model Releases

GTA: Generating Long-Horizon Tasks for Web Agents at Scale

DGX agent

arXiv:2605.29218v1 Announce Type: new Abstract: Web agents, which couple language models with browsing and tool-use capabilities, show promise as open web assistants. Yet progress is increasingly limi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

How LoRA Remembers? A Parametric Memory Law for LLM Finetuning

DGX agent

arXiv:2605.30260v1 Announce Type: cross Abstract: Large Language Models (LLMs) must continuously learn and update knowledge to remain effective in dynamic real-world environments. While Low-Rank Adapt

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Internal Representation, Not Clinical Knowledge: Where Apparent LLM Triage Failures Originate

DGX agent

arXiv:2605.29889v1 Announce Type: cross Abstract: Patient-voiced clinical-triage benchmarks report high under-triage rates for consumer LLMs for constrained multiple-choice output, yet the same cases

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Learning to Extrapolate to New Tasks: A Relational Approach to Task Extrapolation

DGX agent

arXiv:2605.30132v1 Announce Type: new Abstract: Modern learning systems excel at interpolation but struggle to generalize to unseen tasks outside the training distribution's support. This failure occu

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Loong: A Human-Like Long Document Translation Agent with Observe-and-Act Adaptive Context Selection

DGX agent

arXiv:2605.30274v1 Announce Type: cross Abstract: Document-level translation remains one of the most challenging tasks for large language models, which are constrained by limited context windows that

model-releasesarxiv-cs-ai
29 May 2026
Safety

Mask the Target: A Plug-and-Play Regularizer Against LoRA Forgetting

DGX agent

arXiv:2605.29498v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become one of the most widely used fine-tuning mechanisms for adapting large language models to new domains, tasks, and u

safetyarxiv-cs-cl
29 May 2026
Model Releases

Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?

DGX agent

arXiv:2605.28860v1 Announce Type: cross Abstract: Fine-tuning large language models (LLMs) frequently induces catastrophic forgetting of prior capabilities. Recent work has shown that reinforcement le

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval

DGX agent

arXiv:2605.30120v1 Announce Type: cross Abstract: Multi-vector retrieval (MVR) models, exemplified by ColBERT, have established new benchmarks in retrieval accuracy by preserving fine-grained token-le

model-releasesarxiv-cs-ai
29 May 2026
Safety

Offline Reinforcement Learning with Generative Trajectory Policies

DGX agent

arXiv:2510.11499v2 Announce Type: replace-cross Abstract: Generative models have emerged as a powerful class of policies for offline reinforcement learning (RL) due to their ability to capture complex

safetyarxiv-cs-ai
29 May 2026
Model Releases

ParaTool: Shifting Tool Representations from Context to Parameters

DGX agent

arXiv:2605.29561v1 Announce Type: new Abstract: Tool calling extends large language models (LLMs) by enabling grounded interaction with external executable interfaces, thereby supporting environment-c

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Predicting Causal Effects from Natural Language Queries using Structured Representations

DGX agent

arXiv:2605.29631v1 Announce Type: cross Abstract: Randomized controlled trials are a cornerstone of medicine and the social sciences as they enable reliable estimates of causal effects. However, they

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Probabilistic bias adjustment of seasonal forecasts using generative machine learning: A case study of Arctic sea ice predictions

DGX agent

arXiv:2605.29172v1 Announce Type: new Abstract: Seasonal climate predictions support planning and risk management by offering early information of the most likely-to-occur climate conditions in the co

model-releasesarxiv-cs-lg
29 May 2026
Model Releases

Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning

DGX agent

arXiv:2602.00994v2 Announce Type: replace Abstract: Agentic Reinforcement Learning (ARL) trains large language models to interleave reasoning with external tool execution to solve complex tasks. Most

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Rethinking Post-Training Recipes for Multimodal Time-Series Forecasting

DGX agent

arXiv:2605.29401v1 Announce Type: new Abstract: Time-Series Foundation Models (TSFMs) excel at zero-shot unimodal forecasting using numerical data, but unlike LLMs they cannot consume multimodal, non-

model-releasesarxiv-cs-lg
29 May 2026
Tutorials

Robust Cross-Domain Generalization Using Unlabeled Target Data with Source-Domain Supervision

DGX agent

arXiv:2605.29122v1 Announce Type: new Abstract: It is often desirable to generalize medical imaging AI models trained with dense annotations to data acquired from different ultrasound scanners or clin

tutorialsarxiv-cs-cv
29 May 2026
Model Releases

SCDBench: A Benchmark for LLM-Based Smart Contract Decompilers

DGX agent

arXiv:2605.29059v1 Announce Type: cross Abstract: Smart contract decompilation aims to recover high-level source code from bytecode, but evaluating decompilers remains difficult because existing studi

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

SCOPE: Prompt Evolution for Enhancing Agent Effectiveness

DGX agent

arXiv:2512.15374v2 Announce Type: replace Abstract: Large Language Model (LLM) agents are increasingly deployed in environments that generate massive, dynamic contexts. However, a critical bottleneck

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

Uncertainty-Aware Transfer Learning for Cross-Building Energy Forecasting: Toward Robust and Scalable District-Level Energy Management

DGX agent

arXiv:2605.29733v1 Announce Type: new Abstract: Scaling data-driven energy forecasting to district level requires models that can be re-used across buildings with minimal target-domain data and honest

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

When Should a Robot Think? Resource-Aware Reasoning via Reinforcement Learning for Embodied Robotic Decision-Making

DGX agent

arXiv:2603.16673v4 Announce Type: replace-cross Abstract: Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-

model-releasesarxiv-cs-ai
29 May 2026
Model Releases

A Bayesian Nonparametric Perspective on Mahalanobis Distance for Out of Distribution Detection

DGX agent

arXiv:2502.08695v2 Announce Type: replace-cross Abstract: Bayesian nonparametric methods are naturally suited to the problem of out-of-distribution (OOD) detection. However, these techniques have larg

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

A Unified Framework for the Evaluation of LLM Agentic Capabilities

DGX agent

arXiv:2605.27898v1 Announce Type: new Abstract: As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential. However, reported benchmark scores

model-releasesarxiv-cs-ai
28 May 2026
Applications

Architecture-driven Shift: towards a lightweight selector for capturing the trends of logit shift

DGX agent

arXiv:2605.27469v1 Announce Type: cross Abstract: Continual Learning (CL) is a practical paradigm to utilize power of deep pre-trained neural networks, but which pre-trained model has a better ability

applicationsarxiv-cs-ai
28 May 2026
Model Releases

Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity

DGX agent

arXiv:2605.28640v1 Announce Type: new Abstract: Efficient inference is critical for long-context language models, where attention computation and KV-cache access dominate the cost. Recent work RAT+, i

model-releasesarxiv-cs-lg
28 May 2026
Model Releases

Better Accuracies, Worse Reasoning: A Step-Level Audit of Medical Chain-of-Thought Distillation

DGX agent

arXiv:2605.28301v1 Announce Type: new Abstract: Chain-of-thought (CoT) distillation trains a smaller model to imitate a teacher's reasoning trace, but it is typically evaluated by final-answer metrics

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond Motion Primitives: Behavioral Activity Recognition from Head-Mounted IMU

DGX agent

arXiv:2605.27464v1 Announce Type: cross Abstract: AR smart glasses need continuous behavioral context to offer proactive assistance, yet their most practical always-on sensor, the head-mounted Inertia

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

Beyond One Path: Evaluating and Enhancing Divergent Thinking in Interactive LLM Agents

DGX agent

arXiv:2605.28465v1 Announce Type: new Abstract: Divergent thinking is a core dimension of creativity, yet existing evaluations of Large Language Models (LLMs) treat them as single-turn text generation

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Bias Leaves a Gradient Trail: Label-Free Bias Identification via Gradient Probes on Concept Decompositions

DGX agent

arXiv:2605.28780v1 Announce Type: new Abstract: Vision classifiers can exploit spurious correlations, achieving high in-distribution accuracy yet failing under distribution shift. Existing approaches

model-releasesarxiv-cs-cv
28 May 2026
Model Releases

CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text

DGX agent

arXiv:2605.27700v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reports, but they can produce references that appear plausible while contain

model-releasesarxiv-cs-ai
28 May 2026
Model Releases

ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering

DGX agent

arXiv:2605.28093v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) has emerged as a promising paradigm for enhancing large language models (LLMs) on multi-hop question answering (QA)

model-releasesarxiv-cs-cl
28 May 2026
Safety

CPPO: Contrastive Perception Policy Optimization for VLM Agents

DGX agent

arXiv:2601.00501v2 Announce Type: replace Abstract: We introduce CPPO, a Contrastive Perception Policy Optimization method for finetuning vision--language models (VLMs). Reliable perception is a core

safetyarxiv-cs-cv
28 May 2026
← Previous
1…437438439440441…1119
Next →