AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Research

Not Worth Mentioning? A Pilot Study on Salient Proposition Annotation

DGX agent

arXiv:2603.27358v2 Announce Type: replace Abstract: Despite a long tradition of work on extractive summarization, which by nature aims to recover the most important propositions in a text, little work

researcharxiv-cs-cl
13 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Applications

OmniThoughtVis: A Scalable Distillation Pipeline for Deployable Multimodal Reasoning Models

DGX agent

arXiv:2605.11629v1 Announce Type: new Abstract: Recent multimodal large language models (MLLMs) have shown strong chain-of-thought (CoT) reasoning ability on vision-language tasks, but their direct de

applicationsarxiv-cs-cl
13 May 2026
Research

On Predicting the Post-training Potential of Pre-trained LLMs

DGX agent

arXiv:2605.11978v1 Announce Type: new Abstract: The performance of Large Language Models (LLMs) on downstream tasks is fundamentally constrained by the capabilities acquired during pre-training. Howev

researcharxiv-cs-cl
13 May 2026
Agents

On Problems of Implicit Context Compression for Software Engineering Agents

DGX agent

arXiv:2605.11051v1 Announce Type: cross Abstract: LLM-based Software Engineering agents face a critical bottleneck: context length limitations cause failures on complex, long-horizon tasks. One promis

agentsarxiv-cs-cl
13 May 2026
Safety

One Turn Too Late: Response-Aware Defense Against Hidden Malicious Intent in Multi-Turn Dialogue

DGX agent

arXiv:2605.05630v2 Announce Type: replace Abstract: Hidden malicious intent in multi-turn dialogue poses a growing threat to deployed large language models (LLMs). Rather than exposing a harmful objec

safetyarxiv-cs-cl
13 May 2026
Research

ORBIT: Preserving Foundational Language Capabilities in GenRetrieval via Origin-Regulated Merging

DGX agent

arXiv:2605.12419v1 Announce Type: new Abstract: Despite the rapid advancements in large language model (LLM) development, fine-tuning them for specific tasks often results in the catastrophic forgetti

researcharxiv-cs-cl
13 May 2026
Safety

ORCE: Order-Aware Alignment of Verbalized Confidence in Large Language Models

DGX agent

arXiv:2605.12446v1 Announce Type: cross Abstract: Large language models (LLMs) often produce answers with high certainty even when they are incorrect, making reliable confidence estimation essential f

safetyarxiv-cs-cl
13 May 2026
Model Releases

Output Composability of QLoRA PEFT Modules for Plug-and-Play Attribute-Controlled Text Generation

DGX agent

arXiv:2605.12345v1 Announce Type: new Abstract: Parameter-efficient fine-tuning (PEFT) techniques offer task-specific fine-tuning at a fraction of the cost of full fine-tuning, but require separate fi

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Overview of the MedHopQA track at BioCreative IX: track description, participation and evaluation of systems for multi-hop medical question answering

DGX agent

arXiv:2605.12313v1 Announce Type: new Abstract: Multi-hop question answering (QA) remains a significant challenge in the biomedical domain, requiring systems to integrate information across multiple s

model-releasesarxiv-cs-cl
13 May 2026
Agents

Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling

DGX agent

arXiv:2605.12411v1 Announce Type: cross Abstract: AI agents negotiate and transact in natural language with unfamiliar counterparts: a buyer bot facing an unknown seller, or a procurement assistant ne

agentsarxiv-cs-cl
13 May 2026
Research

Predicting Disagreement with Human Raters in LLM-as-a-Judge Difficulty Assessment without Using Generation-Time Probability Signals

DGX agent

arXiv:2605.12422v1 Announce Type: new Abstract: Automatic generation of educational materials using large language models (LLMs) is becoming increasingly common, but assigning difficulty levels to suc

researcharxiv-cs-cl
13 May 2026
Model Releases

Predicting Psychological Well-Being from Spontaneous Speech using LLMs

DGX agent

arXiv:2605.11303v1 Announce Type: new Abstract: We investigate the use of Large Language Models (LLMs) for zero-shot prediction of Ryff Psychological Well-Being (PWB) scores from spontaneous speech. U

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

PreScam: A Benchmark for Predicting Scam Progression from Early Conversations

DGX agent

arXiv:2605.12243v1 Announce Type: new Abstract: Conversational scams, such as romance and investment scams, are emerging as a major form of online fraud. Unlike one-shot scam lures such as fake lotter

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

DGX agent

arXiv:2605.11363v1 Announce Type: cross Abstract: Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal med

model-releasesarxiv-cs-cl
13 May 2026
Safety

Pretraining Exposure Explains Popularity Judgments in Large Language Models

DGX agent

arXiv:2605.12382v1 Announce Type: new Abstract: Large language models (LLMs) exhibit systematic preferences for well-known entities, a phenomenon often attributed to popularity bias. However, the exte

safetyarxiv-cs-cl
13 May 2026
Safety

Primal Generation, Dual Judgment: Self-Training from Test-Time Scaling

DGX agent

arXiv:2605.11299v1 Announce Type: cross Abstract: Code generation is typically trained in the primal space of programs: a model produces a candidate solution and receives sparse execution feedback, of

safetyarxiv-cs-cl
13 May 2026
Local Ai

PRISM: A Geometric Risk Bound that Decomposes Drift into Scale, Shape, and Head

DGX agent

arXiv:2605.11608v1 Announce Type: new Abstract: Comparing post-training LLM variants, such as quantized, LoRA-adapted, and distilled models, requires a diagnostic that identifies how a variant has dri

local-aiarxiv-cs-cl
13 May 2026
Model Releases

PRISM: Pareto-Efficient Retrieval over Intent-Aware Structured Memory for Long-Horizon Agents

DGX agent

arXiv:2605.12260v1 Announce Type: new Abstract: Long-horizon language agents accumulate conversation history far faster than any fixed context window can hold, making memory management critical to bot

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Probabilistic Calibration Is a Trainable Capability in Language Models

DGX agent

arXiv:2605.11845v1 Announce Type: new Abstract: Language models are increasingly used in settings where outputs must satisfy user-specified randomness constraints, yet their generation probabilities a

model-releasesarxiv-cs-cl
13 May 2026
Applications

Prompting from the bench: Large-scale pretraining is not sufficient to prepare LLMs for ordinary meaning analysis

DGX agent

arXiv:2510.25356v2 Announce Type: replace Abstract: In the U.S. judicial system, a widespread approach to legal interpretation entails assessing how a legal text would be understood by an `ordinary' s

applicationsarxiv-cs-cl
13 May 2026
Safety

Question Difficulty Estimation for Large Language Models via Answer Plausibility Scoring

DGX agent

arXiv:2605.12398v1 Announce Type: new Abstract: Estimating question difficulty is a critical component in evaluating and improving large language models (LLMs) for question answering (QA). Existing ap

safetyarxiv-cs-cl
13 May 2026
Model Releases

Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models

DGX agent

arXiv:2605.11887v1 Announce Type: new Abstract: Large language models have achieved remarkable capabilities across diverse tasks, yet their internal decision-making processes remain largely opaque, li

model-releasesarxiv-cs-cl
13 May 2026
Safety

RACC: Representation-Aware Coverage Criteria for LLM Safety Testing

DGX agent

arXiv:2602.02280v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) face severe safety risks from jailbreak attacks, yet current safety testing largely relies on static datasets and

safetyarxiv-cs-cl
13 May 2026
Model Releases

READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling

DGX agent

arXiv:2312.06950v3 Announce Type: replace-cross Abstract: Fully fine-tuning pretrained large-scale transformer models has become a popular paradigm for video-language modeling tasks, such as temporal

model-releasesarxiv-cs-cl
13 May 2026
Research

ReAD: Reinforcement-Guided Capability Distillation for Large Language Models

DGX agent

arXiv:2605.11290v1 Announce Type: new Abstract: Capability distillation applies knowledge distillation to selected model capabilities, aiming to compress a large language model (LLM) into a smaller on

researcharxiv-cs-cl
13 May 2026
Model Releases

Reconstructing Sepsis Trajectories from Clinical Case Reports using LLMs: the Textual Time Series Corpus for Sepsis

DGX agent

arXiv:2504.12326v3 Announce Type: replace Abstract: Clinical case reports and discharge summaries may be the most complete and accurate summarization of patient encounters, yet they are finalized, i.e

model-releasesarxiv-cs-cl
13 May 2026
Applications

Reconstruction of Personally Identifiable Information from Supervised Finetuned Models

DGX agent

arXiv:2605.12264v1 Announce Type: cross Abstract: Supervised Finetuning (SFT) has become one of the primary methods for adapting a large language model (LLM) with extensive pre-trained knowledge to do

applicationsarxiv-cs-cl
13 May 2026
Safety

Red-Teaming Text-to-Image Models via In-Context Experience Replay and Semantic-Preserving Prompt Rewriting

DGX agent

arXiv:2411.16769v3 Announce Type: replace-cross Abstract: Understanding the capabilities of text-to-image (T2I) models in harmful content generation is essential to safety and compliance. However, hum

safetyarxiv-cs-cl
13 May 2026
Tutorials

Reflect then Learn: Active Prompting for Information Extraction Guided by Introspective Confusion

DGX agent

arXiv:2508.10036v2 Announce Type: replace Abstract: Large Language Models (LLMs) show remarkable potential for few-shot information extraction (IE), yet their performance is highly sensitive to the ch

tutorialsarxiv-cs-cl
13 May 2026
Research

RETUYT-INCO at BEA 2026 Shared Task 2: Meta-prompting in Rubric-based Scoring for German

DGX agent

arXiv:2605.11242v1 Announce Type: new Abstract: In this paper, we present the RETUYT-INCO participation at the BEA 2026 shared task 'Rubric-based Short Answer Scoring for German'. Our team participate

researcharxiv-cs-cl
13 May 2026
Research

ReVision: Scaling Computer-Use Agents via Temporal Visual Redundancy Reduction

DGX agent

arXiv:2605.11212v1 Announce Type: new Abstract: Computer-use agents~(CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual toke

researcharxiv-cs-cl
13 May 2026
Research

Robust Biomedical Publication Type and Study Design Classification with Knowledge-Guided Perturbations

DGX agent

arXiv:2605.11502v1 Announce Type: new Abstract: Accurately and consistently indexing biomedical literature by publication type and study design is essential for supporting evidence synthesis and knowl

researcharxiv-cs-cl
13 May 2026
Safety

Robust LLM Unlearning Against Relearning Attacks: The Minor Components in Representations Matter

DGX agent

arXiv:2605.11685v1 Announce Type: new Abstract: Large language model (LLM) unlearning aims to remove specific data influences from pre-trained model without costly retraining, addressing privacy, copy

safetyarxiv-cs-cl
13 May 2026
Model Releases

ROMER: Expert Replacement and Router Calibration for Robust MoE LLMs on Analog Compute-in-Memory Systems

DGX agent

arXiv:2605.11800v1 Announce Type: cross Abstract: Large language models (LLMs) with mixture-of-experts (MoE) architectures achieve remarkable scalability by sparsely activating a subset of experts per

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Routers Learn the Geometry of Their Experts: Geometric Coupling in Sparse Mixture-of-Experts

DGX agent

arXiv:2605.12476v1 Announce Type: cross Abstract: Sparse Mixture-of-Experts (SMoE) models enable scaling language models efficiently, but training them remains challenging, as routing can collapse ont

model-releasesarxiv-cs-cl
13 May 2026
Safety

Safety Alignment as Continual Learning: Mitigating the Alignment Tax via Orthogonal Gradient Projection

DGX agent

arXiv:2602.07892v2 Announce Type: replace-cross Abstract: Safety post-training can improve the harmfulness and policy compliance of Large Language Models (LLMs), but it may also reduce general utility

safetyarxiv-cs-cl
13 May 2026
Safety

Safety-Oriented Evaluation of Language Understanding Systems for Air Traffic Control

DGX agent

arXiv:2605.11769v1 Announce Type: new Abstract: Air Traffic Control (ATC) is a safety-critical domain in which incorrect interpretation of instructions may lead to severe operational consequences. Whi

safetyarxiv-cs-cl
13 May 2026
Model Releases

SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation

DGX agent

arXiv:2605.12022v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong performance on standard knowledge evaluation benchmarks, yet recent work shows that their knowledge capabili

model-releasesarxiv-cs-cl
13 May 2026
Local Ai

Sampling More, Getting Less: Calibration is the Diversity Bottleneck in LLMs

DGX agent

arXiv:2605.11128v1 Announce Type: new Abstract: Diversity is essential for language-model applications ranging from creative generation to scientific discovery, yet modern LLMs often collapse into a n

local-aiarxiv-cs-cl
13 May 2026
Research

Scalable Token-Level Hallucination Detection in Large Language Models

DGX agent

arXiv:2605.12384v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated remarkable capabilities, but they still frequently produce hallucinations. These hallucinations are diffi

researcharxiv-cs-cl
13 May 2026
Research

Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models

DGX agent

arXiv:2605.11854v1 Announce Type: new Abstract: Diffusion Language Models (DLMs) have recently emerged as a promising alternative to autoregressive language models, offering stronger global awareness

researcharxiv-cs-cl
13 May 2026
Applications

Sign Language Recognition and Translation for Low-Resource Languages: Challenges and Pathways Forward

DGX agent

arXiv:2605.12096v1 Announce Type: new Abstract: Sign languages are natural, visual-gestural languages used by Deaf communities worldwide. Over 300 distinct sign languages remain severely low-resource

applicationsarxiv-cs-cl
13 May 2026
Safety

SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs

DGX agent

arXiv:2605.12039v1 Announce Type: new Abstract: Skill libraries enable large language model agents to reuse experience from past interactions, but most existing libraries store skills as isolated entr

safetyarxiv-cs-cl
13 May 2026
Model Releases

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

DGX agent

arXiv:2605.12015v1 Announce Type: cross Abstract: Reusable skills are becoming a common interface for extending large language model agents, packaging procedural guidance with access to files, tools,

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Slicing and Dicing: Configuring Optimal Mixtures of Experts

DGX agent

arXiv:2605.11689v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures have become standard in large language models, yet many of their core design choices - expert count, granularit

model-releasesarxiv-cs-cl
13 May 2026
Model Releases

Solve the Loop: Attractor Models for Language and Reasoning

DGX agent

arXiv:2605.12466v1 Announce Type: cross Abstract: Looped Transformers offer a promising alternative to purely feed-forward computation by iteratively refining latent representations, improving languag

model-releasesarxiv-cs-cl
13 May 2026
Local Ai

SOMA: Efficient Multi-turn LLM Serving via Small Language Model

DGX agent

arXiv:2605.11317v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly deployed in multi-turn dialogue settings where preserving conversational context across turns is essential

local-aiarxiv-cs-cl
13 May 2026
Research

Sparse Attention Remapping with Clustering for Efficient LLM Decoding on PIM

DGX agent

arXiv:2505.05772v2 Announce Type: replace Abstract: Transformer-based models are the foundation of modern machine learning, but their execution, particularly during autoregressive decoding in large la

researcharxiv-cs-cl
13 May 2026
← Previous
1…9596979899…161
Next →