AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
13 May 2026

PreScam: A Benchmark for Predicting Scam Progression from Early Conversations

Model ReleasesDGX agent

arXiv:2605.12243v1 Announce Type: new Abstract: Conversational scams, such as romance and investment scams, are emerging as a major form of online fraud. Unlike one-shot scam lures such as fake lotter

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

Model ReleasesDGX agent

arXiv:2605.11363v1 Announce Type: cross Abstract: Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal med

Pretraining Exposure Explains Popularity Judgments in Large Language Models

SafetyDGX agent

arXiv:2605.12382v1 Announce Type: new Abstract: Large language models (LLMs) exhibit systematic preferences for well-known entities, a phenomenon often attributed to popularity bias. However, the exte


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling

SafetyDGX agent

arXiv:2605.11299v1 Announce Type: cross Abstract: Code generation is typically trained in the primal space of programs: a model produces a candidate solution and receives sparse execution feedback, of

PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head

Local AiDGX agent

arXiv:2605.11608v1 Announce Type: new Abstract: Comparing post-training LLM variants, such as quantized, LoRA-adapted, and distilled models, requires a diagnostic that identifies how a variant has dri

PRISM: Pareto-Efficient Retrieval over Intent-Aware Structured Memory for Long-Horizon Agents

Model ReleasesDGX agent

arXiv:2605.12260v1 Announce Type: new Abstract: Long-horizon language agents accumulate conversation history far faster than any fixed context window can hold, making memory management critical to bot

Probabilistic Calibration Is a Trainable Capability in Language Models

Model ReleasesDGX agent

arXiv:2605.11845v1 Announce Type: new Abstract: Language models are increasingly used in settings where outputs must satisfy user-specified randomness constraints, yet their generation probabilities a

Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis

ApplicationsDGX agent

arXiv:2510.25356v2 Announce Type: replace Abstract: In the U.S. judicial system, a widespread approach to legal interpretation entails assessing how a legal text would be understood by an `ordinary' s

Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring

SafetyDGX agent

arXiv:2605.12398v1 Announce Type: new Abstract: Estimating question difficulty is a critical component in evaluating and improving large language models (LLMs) for question answering (QA). Existing ap

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

Model ReleasesDGX agent

arXiv:2605.11887v1 Announce Type: new Abstract: Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, li

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing

SafetyDGX agent

arXiv:2602.02280v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face severe safety risks from jailbreak attacks, yet current safety testing largely relies on static datasets and

READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling

Model ReleasesDGX agent

arXiv:2312.06950v3 Announce Type: replace-cross Abstract: Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal

ReAD: Reinforcement-Guided Capability Distillation for Large Language Models

ResearchDGX agent

arXiv:2605.11290v1 Announce Type: new Abstract: Capability distillation applies knowledge distillation to selected model capabilities, aiming to compress a large language model (LLM) into a smaller on

Reconstructing Sepsis Trajectories from Clinical Case Reports using LLMs: the Textual Time Series Corpus for Sepsis

Model ReleasesDGX agent

arXiv:2504.12326v3 Announce Type: replace Abstract: Clinical case reports and discharge summaries may be the most complete and accurate summarization of patient encounters, yet they are finalized, i.e

Reconstruction of Personally Identifiable Information from Supervised Finetuned Models

ApplicationsDGX agent

arXiv:2605.12264v1 Announce Type: cross Abstract: Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to do

Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

SafetyDGX agent

arXiv:2411.16769v3 Announce Type: replace-cross Abstract: Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, hum

Reflect then Learn: Active Prompting for Information Extraction Guided by Introspective Confusion

TutorialsDGX agent

arXiv:2508.10036v2 Announce Type: replace Abstract: Large Language Models (LLMs) show remarkable potential for few-shot information extraction (IE), yet their performance is highly sensitive to the ch

RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German

ResearchDGX agent

arXiv:2605.11242v1 Announce Type: new Abstract: In this paper, we present the RETUYT-INCO participation at the BEA 2026 shared task 'Rubric-based Short Answer Scoring for German'. Our team participate

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction

ResearchDGX agent

arXiv:2605.11212v1 Announce Type: new Abstract: Computer-use agents~(CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual toke

Robust Biomedical Publication Type and Study Design Classification with Knowledge-Guided Perturbations

ResearchDGX agent

arXiv:2605.11502v1 Announce Type: new Abstract: Accurately and consistently indexing biomedical literature by publication type and study design is essential for supporting evidence synthesis and knowl

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

SafetyDGX agent

arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copy

ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems

Model ReleasesDGX agent

arXiv:2605.11800v1 Announce Type: cross Abstract: Large language models (LLMs) with mixture-of-experts (MoE) architectures achieve remarkable scalability by sparsely activating a subset of experts per

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts

Model ReleasesDGX agent

arXiv:2605.12476v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse ont

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

SafetyDGX agent

arXiv:2602.07892v2 Announce Type: replace-cross Abstract: Safety post-training can improve the harmfulness and policy compliance of Large Language Models (LLMs), but it may also reduce general utility

Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control

SafetyDGX agent

arXiv:2605.11769v1 Announce Type: new Abstract: Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. Whi

SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation

Model ReleasesDGX agent

arXiv:2605.12022v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabili

Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs

Local AiDGX agent

arXiv:2605.11128v1 Announce Type: new Abstract: Diversity is essential for language-model applications ranging from creative generation to scientific discovery, yet modern LLMs often collapse into a n

Scalable Token-Level Hallucination Detection in Large Language Models

ResearchDGX agent

arXiv:2605.12384v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities, but they still frequently produce hallucinations. These hallucinations are diffi

Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models

ResearchDGX agent

arXiv:2605.11854v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive language models, offering stronger global awareness

Sign Language Recognition and Translation for Low-Resource Languages: Challenges and Pathways Forward

ApplicationsDGX agent

arXiv:2605.12096v1 Announce Type: new Abstract: Sign languages are natural, visual-gestural languages used by Deaf communities worldwide. Over 300 distinct sign languages remain severely low-resource

SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs

SafetyDGX agent

arXiv:2605.12039v1 Announce Type: new Abstract: Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entr

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

Model ReleasesDGX agent

arXiv:2605.12015v1 Announce Type: cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools,

Slicing and Dicing: Configuring Optimal Mixtures of Experts

Model ReleasesDGX agent

arXiv:2605.11689v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become standard in large language models, yet many of their core design choices - expert count, granularit

Solve the Loop: Attractor Models for Language and Reasoning

Model ReleasesDGX agent

arXiv:2605.12466v1 Announce Type: cross Abstract: Looped Transformers offer a promising alternative to purely feed-forward computation by iteratively refining latent representations, improving languag

SOMA: Efficient Multi-turn LLM Serving via Small Language Model

Local AiDGX agent

arXiv:2605.11317v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in multi-turn dialogue settings where preserving conversational context across turns is essential

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM

ResearchDGX agent

arXiv:2505.05772v2 Announce Type: replace Abstract: Transformer-based models are the foundation of modern machine learning, but their execution, particularly during autoregressive decoding in large la

STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens

Model ReleasesDGX agent

arXiv:2602.15620v4 Announce Type: replace Abstract: Reinforcement Learning (RL) has significantly improved large language model reasoning, but existing RL fine-tuning methods rely heavily on heuristic

Steering Without Breaking: Mechanistically Informed Interventions for Discrete Diffusion Language Models

ResearchDGX agent

arXiv:2605.10971v1 Announce Type: cross Abstract: Discrete diffusion language models (DLMs) generate text by iteratively denoising all positions in parallel, offering an alternative to autoregressive

StepCodeReasoner: Aligning Code Reasoning with Stepwise Execution Traces via Reinforcement Learning

Model ReleasesDGX agent

arXiv:2605.11922v1 Announce Type: cross Abstract: Existing code reasoning methods primarily supervise final code outputs, ignoring intermediate states, often leading to reward hacking where correct an

StoicLLM: Preference Optimization for Philosophical Alignment in Small Language Models

Model ReleasesDGX agent

arXiv:2605.11483v1 Announce Type: new Abstract: While large language models excel at factual adaptation, their ability to internalize nuanced philosophical frameworks under severe data constraints rem

Stopping Computation for Converged Tokens in Masked Diffusion-LM Decoding

ResearchDGX agent

arXiv:2602.06412v3 Announce Type: replace Abstract: Masked Diffusion Language Models generate sequences via iterative sampling that progressively unmasks tokens. However, they still recompute the atte

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space

ResearchDGX agent

arXiv:2605.12412v1 Announce Type: new Abstract: Large Language Models (LLMs) update their behavior in context, which can be viewed as a form of Bayesian inference. However, the structure of the latent

Synthetic Function Demonstrations Improve Generation in Low-Resource Programming Languages

TutorialsDGX agent

arXiv:2503.18760v2 Announce Type: replace Abstract: A key consideration when training an LLM is whether the target language is more or less resourced, for example English compared to Welsh, or Python

TabDLM: Free-Form Tabular Data Generation via Joint Numerical-Language Diffusion

ApplicationsDGX agent

arXiv:2602.22586v2 Announce Type: replace-cross Abstract: Synthetic tabular data generation has attracted growing attention due to its importance for data augmentation, foundation models, and privacy.

Taming Extreme Tokens: Covariance-Aware GRPO with Gaussian-Kernel Advantage Reweighting

SafetyDGX agent

arXiv:2605.11538v1 Announce Type: new Abstract: Group Relative Policy Optimization (GRPO) has emerged as a promising approach for improving the reasoning capabilities of large language models. However

Task-Adaptive Embedding Refinement via Test-time LLM Guidance

ResearchDGX agent

arXiv:2605.12487v1 Announce Type: new Abstract: We explore the effectiveness of an LLM-guided query refinement paradigm for extending the usability of embedding models to challenging zero-shot search

Test-Time Compute for Dense Retrieval: Agentic Program Generation with Frozen Embedding Models

Model ReleasesDGX agent

arXiv:2605.11374v1 Announce Type: cross Abstract: Test-time compute is widely believed to benefit only large reasoning models. We show it also helps small embedding models. Most modern embedding check

TextSeal: A Localized LLM Watermark for Provenance & Distillation Protection

Local AiDGX agent

arXiv:2605.12456v1 Announce Type: cross Abstract: We introduce TextSeal, a state-of-the-art watermark for large language models. Building on Gumbel-max sampling, TextSeal introduces dual-key generatio

The Algorithmic Caricature: Auditing LLM-Generated Political Discourse Across Crisis Events

ResearchDGX agent

arXiv:2605.12452v1 Announce Type: new Abstract: Large Language Models (LLMs) can generate fluent political text at scale, raising concerns about synthetic discourse during crises and social conflict.

The Bicameral Model: Bidirectional Hidden-State Coupling Between Parallel Language Models

ResearchDGX agent

arXiv:2605.11167v1 Announce Type: new Abstract: Existing multi-model and tool-augmented systems communicate by generating text, serializing every exchange through the output vocabulary. Can two pretra

The Challenge and Reward of Fair Play in Narrative: A Computational Approach

ResearchDGX agent

arXiv:2507.13841v2 Announce Type: replace Abstract: Good storytelling involves surprise -- unpredictability in how the story unfolds -- and sense-making, the requirement that the story forms a coheren

Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation

Model ReleasesDGX agent

arXiv:2605.11574v1 Announce Type: new Abstract: The literature on how large language models handle conflict between their training knowledge and a contradicting document presents a persistent empirica

To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation

HardwareDGX agent

arXiv:2412.14461v4 Announce Type: replace Abstract: Unstructured text data annotation is foundational to management research. LLMs offer a cost-effective and scalable alternative to human annotation,

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching

SafetyDGX agent

arXiv:2605.12288v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) is a widely used RL-free method for aligning language models from pairwise preferences, but it models preferences o

Towards Fine-Grained Code-Switch Speech Translation with Semantic Space Alignment

SafetyDGX agent

arXiv:2511.10670v2 Announce Type: replace Abstract: Code-switching (CS) speech translation (ST) aims to translate speech that alternates between multiple languages into a target language text, posing

Towards Visually-Guided Movie Subtitle Translation for Indic Languages

ApplicationsDGX agent

arXiv:2605.11993v1 Announce Type: new Abstract: Movie subtitle translation is inherently multimodal, yet text-only systems often miss visual cues needed to convey emotion, action, and social nuance, e

Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness

SafetyDGX agent

arXiv:2503.16072v4 Announce Type: replace-cross Abstract: Toxicity detection has become core safety infrastructure for online moderation, dataset filtering, and deployed language-model systems. Yet mo

Training-Inference Consistent Segmented Execution for Long-Context LLMs

ResearchDGX agent

arXiv:2605.11744v1 Announce Type: new Abstract: Transformer-based large language models face severe scalability challenges in long-context generation due to the computational and memory costs of full-

Understanding the Performance Gap in Preference Learning: A Dichotomy of RLHF and DPO

SafetyDGX agent

arXiv:2505.19770v5 Announce Type: replace-cross Abstract: We present a fine-grained theoretical analysis of the performance gap between two-stage reinforcement learning from human feedback~(RLHF) and

UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs

ApplicationsDGX agent

arXiv:2605.11856v1 Announce Type: cross Abstract: Multimodal large language models are increasingly expected to perform thinking with images, yet existing visual latent reasoning methods still rely on

← Previous
1…7677787980…129
Next →