AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,433
  • Agents7,256
  • Applications5,196
  • Concepts5
  • Hardware1,747
  • Industry6,090
  • Local Ai4,704
  • Model Releases22,499
  • Research19,191
  • Safety12,806
  • Syntheses17
  • Tools1,665
  • Tutorials3,257

Source
HumanDGX agent

Content type
AllBlog
84,433Total entries
1Added by human
84,432Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,759 results
Safety

Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction

DGX agent

arXiv:2605.27878v1 Announce Type: new Abstract: Large language models produce fluent fiction, yet their creative output is widely seen as flat. We ask where this quality originates in the training and

safetyarxiv-cs-cl
28 May 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

On Compositional Learning Behaviours in Formal Mathematics

DGX agent

arXiv:2605.28512v1 Announce Type: new Abstract: Self-evolving scientific agents capable of conquering the hard tail of formal mathematics require Compositional Learning Behaviours (CLBs) -- the capaci

model-releasesarxiv-cs-cl
28 May 2026
Applications

OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models

DGX agent

arXiv:2605.27916v1 Announce Type: cross Abstract: The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to suppor

applicationsarxiv-cs-cl
28 May 2026
Model Releases

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

DGX agent

arXiv:2605.27378v1 Announce Type: new Abstract: Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have pro

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI

DGX agent

arXiv:2605.27545v1 Announce Type: new Abstract: Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation

DGX agent

arXiv:2601.18006v2 Announce Type: replace Abstract: We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-fre

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective

DGX agent

arXiv:2605.28819v1 Announce Type: cross Abstract: Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstrea

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

Personal Visual Memory from Explicit and Implicit Evidence

DGX agent

arXiv:2605.28806v1 Announce Type: cross Abstract: Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when imag

model-releasesarxiv-cs-cl
28 May 2026
Agents

Personality, Role, and Expressive Style in Large Language Models: An Interactionist Analysis

DGX agent

arXiv:2605.28037v1 Announce Type: new Abstract: Prompt-based personality control is a key technique for designing large language model (LLM) dialogue agents that behave consistently across social cont

agentsarxiv-cs-cl
28 May 2026
Research

Playing with Words, Improving with Rewards: Training Language Models for Creative Association

DGX agent

arXiv:2605.27832v1 Announce Type: new Abstract: Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLM

researcharxiv-cs-cl
28 May 2026
Model Releases

PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature

DGX agent

arXiv:2605.28375v1 Announce Type: new Abstract: Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stage

model-releasesarxiv-cs-cl
28 May 2026
Safety

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

DGX agent

arXiv:2510.06974v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases

safetyarxiv-cs-cl
28 May 2026
Model Releases

Prompting Is All You Need: Multi-view Prompting Large Language Models for Aspect-Based Sentiment Analysis

DGX agent

arXiv:2605.28058v1 Announce Type: new Abstract: Recent work explored the capabilities of Large Language Models (LLMs) in Aspect-Based Sentiment Analysis (ABSA) through few-shot prompting, requiring su

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text

DGX agent

arXiv:2605.28363v1 Announce Type: new Abstract: Causal relation extraction (CRE) is central to biomedical text mining, but current resources often conflate causal relations with broader associations,

model-releasesarxiv-cs-cl
28 May 2026
Safety

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

DGX agent

arXiv:2602.15894v2 Announce Type: replace Abstract: In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, exist

safetyarxiv-cs-cl
28 May 2026
Agents

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

DGX agent

arXiv:2605.28003v1 Announce Type: new Abstract: The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully en

agentsarxiv-cs-cl
28 May 2026
Research

Rethinking Visual Neglect: Steering via Context-Preference for MLLM Hallucination Mitigation

DGX agent

arXiv:2605.27993v1 Announce Type: new Abstract: Object hallucination remains a primary obstacle to the reliable deployment of Multimodal Large Language Models (MLLMs). Current inference-time mitigatio

researcharxiv-cs-cl
28 May 2026
Safety

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

DGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

safetyarxiv-cs-cl
28 May 2026
Research

ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation

DGX agent

arXiv:2605.27709v1 Announce Type: new Abstract: Mathematical reasoning benchmarks are vital for evaluating large language models (LLMs), but many are static and repeatedly exposed through public evalu

researcharxiv-cs-cl
28 May 2026
Research

Risk-aware Selective Prompting for Hallucination Mitigation in Large Vision-Language Models

DGX agent

arXiv:2605.28123v1 Announce Type: new Abstract: Prompt-based verification is widely used to mitigate hallucinations in large vision-language models (LVLMs), yet when it helps remains poorly understood

researcharxiv-cs-cl
28 May 2026
Model Releases

RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction

DGX agent

arXiv:2602.13748v2 Announce Type: replace Abstract: Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires groundin

model-releasesarxiv-cs-cl
28 May 2026
Agents

Roles with Rails: Contract-Preserving Role Evolution in Multi-Agent Structured Reasoning

DGX agent

arXiv:2605.28433v1 Announce Type: new Abstract: Role-based LLM multi-agent systems need adaptive role pools, yet adapting such systems is not merely a matter of prompt optimization: roles often carry

agentsarxiv-cs-cl
28 May 2026
Safety

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

DGX agent

arXiv:2605.28014v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-

safetyarxiv-cs-cl
28 May 2026
Research

Self-Consistency via Marginal Sharpening

DGX agent

arXiv:2605.28142v1 Announce Type: cross Abstract: Inference-time sampling can elicit strong reasoning abilities from language models without additional training. Existing power-sampling methods do so

researcharxiv-cs-cl
28 May 2026
Agents

Self-Improving Language Models with Bidirectional Evolutionary Search

DGX agent

arXiv:2605.28814v1 Announce Type: new Abstract: Search has been proposed as an effective method for self-improving language models and agentic systems, both for post-training sample generation and for

agentsarxiv-cs-cl
28 May 2026
Local Ai

Sentence Curve Language Models

DGX agent

arXiv:2602.01807v3 Announce Type: replace Abstract: Language models (LMs) are a central component of modern AI systems, and diffusion language models (DLMs) have recently emerged as a competitive alte

local-aiarxiv-cs-cl
28 May 2026
Research

SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adversarial Data Poisoning

DGX agent

arXiv:2605.28074v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) mitigates LLM hallucinations but introduces a critical vulnerability: corpus integrity. We present SilentRetrieva

researcharxiv-cs-cl
28 May 2026
Model Releases

Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering

DGX agent

arXiv:2605.27636v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate excellent capabilities and performance for general reasoning tasks within the general public domain, t

model-releasesarxiv-cs-cl
28 May 2026
Agents

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

DGX agent

arXiv:2605.27955v1 Announce Type: cross Abstract: Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation synt

agentsarxiv-cs-cl
28 May 2026
Agents

Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning

DGX agent

arXiv:2605.28424v1 Announce Type: new Abstract: Equipping large language models with explicit skills has emerged as a promising paradigm for enabling autonomous agents to solve complex tasks. Agent sk

agentsarxiv-cs-cl
28 May 2026
Safety

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

DGX agent

arXiv:2605.28561v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be che

safetyarxiv-cs-cl
28 May 2026
Research

Stance Detection in Prediction Markets: Addressing Imbalanced Trader Commentary via Counterfactual Augmentation and Market Context

DGX agent

arXiv:2605.28745v1 Announce Type: new Abstract: Prediction markets such as Polymarket aggregate crowd beliefs into real-time probability estimates, and the comments traders post beneath each market co

researcharxiv-cs-cl
28 May 2026
Research

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

DGX agent

arXiv:2602.05897v2 Announce Type: replace Abstract: As large language models become smaller and more efficient, small reasoning models (SRMs) are crucial for enabling chain-of-thought (CoT) reasoning

researcharxiv-cs-cl
28 May 2026
Model Releases

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

DGX agent

arXiv:2605.28179v1 Announce Type: new Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream b

model-releasesarxiv-cs-cl
28 May 2026
Safety

Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect

DGX agent

arXiv:2605.28225v1 Announce Type: new Abstract: Cross-cultural comparison of psychological meaning requires methods that go beyond word-level translation and examine how semantic dimensions are organi

safetyarxiv-cs-cl
28 May 2026
Research

Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese

DGX agent

arXiv:2605.27874v1 Announce Type: new Abstract: Most Automatic Speech Recognition (ASR) systems formulate transcription as a prediction problem over orthographic units such as characters, subwords, or

researcharxiv-cs-cl
28 May 2026
Research

TARQ: Tail-Aware Reconstruction Quantization for Rare-Word Robust Automatic Speech Recognition

DGX agent

arXiv:2605.27808v1 Announce Type: new Abstract: Data-aware post-training quantization (PTQ) minimizes a per-token reconstruction loss on a small calibration corpus, implicitly weighting positions by t

researcharxiv-cs-cl
28 May 2026
Model Releases

The Abstraction Gap in Vision-Language Causal Reasoning

DGX agent

arXiv:2605.28779v1 Announce Type: new Abstract: Vision-language models (VLMs) generate fluent causal explanations, but current evaluations cannot distinguish linguistic plausibility from faithful caus

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

The Harder Text Embedding Benchmark (HTEB): Beyond One-dimensional Static Robustness

DGX agent

arXiv:2605.28190v1 Announce Type: new Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embedding rob

model-releasesarxiv-cs-cl
28 May 2026
Model Releases

The Missing Piece in Pre-trained Model Evaluation: Reward-Guided Decoding Unlocks Task-Oriented Behavior Without Parameter Updates

DGX agent

arXiv:2605.28020v1 Announce Type: new Abstract: With the rapid progress of large language models (LLMs), reliably evaluating the capabilities of pre-trained LLMs has become increasingly important. The

model-releasesarxiv-cs-cl
28 May 2026
Safety

TRACES: Proactive Safety Auditing for Multi-Turn LLM Agents via Trajectory-State Modeling

DGX agent

arXiv:2605.27690v1 Announce Type: new Abstract: LLM agents increasingly operate through multi-turn tool use and environment interaction, where safety risks often emerge from intermediate steps long be

safetyarxiv-cs-cl
28 May 2026
Safety

Trust Me, I'm an Expert: Decoding and Steering Authority Bias in Large Language Models

DGX agent

arXiv:2601.13433v3 Announce Type: replace Abstract: Prior research demonstrates that performance of language models on reasoning tasks can be influenced by suggestions, hints and endorsements. However

safetyarxiv-cs-cl
28 May 2026
Research

UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training

DGX agent

arXiv:2605.27740v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the linear growth of the self-attention key-value (KV) cache. Top-k sparse att

researcharxiv-cs-cl
28 May 2026
Research

ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention

DGX agent

arXiv:2602.07574v2 Announce Type: replace-cross Abstract: Modern multimodal large language models (MLLMs) adopt a unified self-attention design that processes visual and textual tokens at every Transf

researcharxiv-cs-cl
28 May 2026
Safety

VLMs May Not Globally Enhance Human Alignment over LLMs During Natural Reading

DGX agent

arXiv:2605.28818v1 Announce Type: new Abstract: Large language models (LLMs) have become increasingly useful computational models of human language processing, but it remains unclear whether vision-la

safetyarxiv-cs-cl
28 May 2026
Safety

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

DGX agent

arXiv:2601.07648v2 Announce Type: replace Abstract: As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) ad

safetyarxiv-cs-cl
28 May 2026
Research

When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models

DGX agent

arXiv:2605.28181v1 Announce Type: new Abstract: Diffusion language models decode text by iteratively denoising masked token sequences, making the choice of which positions to decode a central inferenc

researcharxiv-cs-cl
28 May 2026
Research

When Discourse Pressures Conflict: Information Structure in Vision-Language Model Outputs

DGX agent

arXiv:2605.28346v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly evaluated for whether they identify the right visual content, but little is known about whether they expr

researcharxiv-cs-cl
28 May 2026
← Previous
1…7273747576…162
Next →