AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
28 May 2026

GraphLit: Learning Text-Enriched Dynamic Character Network Representations for Literary Study

ResearchDGX agent

arXiv:2605.28643v1 Announce Type: new Abstract: Methods to represent literary texts as graphs or sequences of graphs mainly focus on representing character interactions, and often overlook another cru

GraphSteal: Structural Knowledge Stealing from Graph RAG via Traversal Reconstruction

ApplicationsDGX agent

arXiv:2605.28645v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) enhances LLMs by grounding generation in query-relevant external evidence. Beyond unstructured text corpora, Grap

GUI-CIDER: Mid-training GUI Agents via Causal Internalization and Density-aware Exemplar Reselection

AgentsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2605.28534v1 Announce Type: new Abstract: Despite the rapid progress of multimodal large language models in building Graphical User Interface (GUI) agents, their real-world task completion is fu

HardMTBench: Stress-Testing Chinese-English Translation on Knowledge-Intensive Domains

Model ReleasesDGX agent

arXiv:2605.28315v1 Announce Type: new Abstract: General-purpose machine translation benchmarks such as FLORES-200 have reached a saturation regime on Chinese-English pairs, where modern large language

HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment

Model ReleasesDGX agent

arXiv:2605.28308v1 Announce Type: new Abstract: Entity Alignment (EA) is essential for knowledge graph (KG) fusion, but existing benchmarks often allow models to exploit name overlap rather than relat

Human Label Variation as Stable Signal: Learning Annotator-Specific Explanation Behavior via Cross-Annotator Preference Optimization

TutorialsDGX agent

arXiv:2605.28802v1 Announce Type: new Abstract: Free-text explanations extend human label variation (HLV) beyond label disagreement by revealing the reasoning and preferences behind annotators' decisi

ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment

SafetyDGX agent

arXiv:2605.27374v1 Announce Type: new Abstract: Recent advances in multimodal large language models (MLLMs) and diffusion models (DMs) have opened new possibilities for AI-generated content. Yet, pers

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

SafetyDGX agent

arXiv:2601.04716v3 Announce Type: replace Abstract: While Large Language Model (LLM) role-playing agents have advanced rapidly, it remains unclear which profile elements genuinely drive role-playing q

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following

Model ReleasesDGX agent

arXiv:2605.28218v1 Announce Type: new Abstract: Modern translation workflows demand more than semantic equivalence. Users routinely require models to preserve JSON or HTML schemas, honor curated gloss

Interpretability-Guided Layer Selection over Subspace Projection: SAEs as Stethoscopes, Not Scalpels, for Raw Task Vector Model Editing

Model ReleasesDGX agent

arXiv:2605.28649v1 Announce Type: cross Abstract: LLMs increasingly require surgical model editing to enhance domain-specific capabilities without incurring the computational cost or catastrophic forg

Keyphrase Generative Representation of Youth Crisis Conversations Beyond Static Taxonomies

TutorialsDGX agent

arXiv:2605.27546v1 Announce Type: new Abstract: Crisis Responders (CRs) rapidly assess thousands of youth SMS conversations each year to identify mental health concerns and guide support. Yet youth di

Knowing When to Ask: Segment-Level Credit Assignment for LLM Tool Use

AgentsDGX agent

arXiv:2605.27788v1 Announce Type: cross Abstract: Humans know when to reach for help e.g. 347 imes 28 warrants a calculator while 2+2 does not. Language models do not. Prompt-based approaches can inst

Knowledge Dependency Estimation for Reliable Question Answering

ResearchDGX agent

arXiv:2605.28047v1 Announce Type: new Abstract: Reliable question answering requires identifying not only whether an answer is correct, but also which available knowledge the prediction depends on. In

KSAFE-MM: A Multimodal Safety Benchmark via Localized Contextualization for Korean Cultural Risks

Model ReleasesDGX agent

arXiv:2605.28013v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) exacerbate safety risks by introducing vulnerabilities across multiple modalities, such as language and vision.

Large Language Models Approach Expert Pedagogical Quality in Math Tutoring but Differ in Instructional and Linguistic Profiles

AgentsDGX agent

arXiv:2512.20780v3 Announce Type: replace Abstract: Recent work has explored the use of large language models (LLMs) to generate tutoring responses in mathematics, yet it remains unclear how closely t

Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis

AgentsDGX agent

arXiv:2601.16800v3 Announce Type: replace Abstract: Fine-grained opinion analysis of text provides a detailed understanding of expressed sentiments, including the addressed entity. Although this level

LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks

Model ReleasesDGX agent

arXiv:2605.27375v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context

Learning Deliberately, Acting Intuitively: Unlocking Test-Time Reasoning in Multimodal LLMs

SafetyDGX agent

arXiv:2507.06999v2 Announce Type: replace-cross Abstract: Reasoning is essential for large language models (LLMs), especially in complex tasks such as mathematical problem solving. However, multimodal

Learning to Translate from Soft to Hard LLM Prompts

Model ReleasesDGX agent

arXiv:2605.27642v1 Announce Type: new Abstract: Soft prompt tuning is a parameter-efficient method for adapting LLMs to specific tasks, but suffers from a lack of interpretability. Building on recent

Long Live the Librarian! A Persistent Search Sub-Agent for Energy-Efficient Multi-Agent Software Engineering Systems

HardwareDGX agent

arXiv:2605.27787v1 Announce Type: cross Abstract: Multi-agent systems (MAS) have substantially advanced autonomous software engineering (SWE), but their growing inference energy demands raise sustaina

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

Model ReleasesDGX agent

arXiv:2603.21165v2 Announce Type: replace Abstract: Bangla culture is richly expressed through region, dialect, history, food, politics, media, and everyday visual life, yet it remains underrepresente

MaskClaw: Edge-Side Personalized Privacy Arbitration for GUI Agents with Behavior-Driven Skill Evolution

Model ReleasesDGX agent

arXiv:2605.28646v1 Announce Type: cross Abstract: GUI agents rely on screenshots to infer intent and operate across applications, but these screenshots often contain private messages, medical records,

MERIT: Matching Expertise via Rubric-Informed Training for Reviewer Assignment

ResearchDGX agent

arXiv:2605.27865v1 Announce Type: new Abstract: Matching submissions with suitable reviewers at scale is a growing challenge for major venues, yet existing approaches either rely on coarse proxy signa

Mitigating Cross-Lingual Cultural Inconsistencies in LLMs via Consensus-Driven Preference Optimisation

SafetyDGX agent

arXiv:2605.12515v2 Announce Type: replace Abstract: Despite their impressive capabilities, multilingual large language models (MLLMs) frequently exhibit inconsistent behaviour when the prompt's langua

Mobile-Aptus: Confidence-Driven Proactive and Robust Interaction in MLLM-based Mobile-Using Agents

SafetyDGX agent

arXiv:2605.28629v1 Announce Type: new Abstract: Recent advancements in multimodal large language models (MLLMs) have shown exceptional potential in enabling mobile-using agents to autonomously execute

Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction

SafetyDGX agent

arXiv:2605.27878v1 Announce Type: new Abstract: Large language models produce fluent fiction, yet their creative output is widely seen as flat. We ask where this quality originates in the training and

On Compositional Learning Behaviours in Formal Mathematics

Model ReleasesDGX agent

arXiv:2605.28512v1 Announce Type: new Abstract: Self-evolving scientific agents capable of conquering the hard tail of formal mathematics require Compositional Learning Behaviours (CLBs) -- the capaci

OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models

ApplicationsDGX agent

arXiv:2605.27916v1 Announce Type: cross Abstract: The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to suppor

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

Model ReleasesDGX agent

arXiv:2605.27378v1 Announce Type: new Abstract: Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have pro

PAST2HARM: A Simple Adaptive Past Tense Attack for Jailbreaking Multimodal AI

Model ReleasesDGX agent

arXiv:2605.27545v1 Announce Type: new Abstract: Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text

PEAR: Pairwise Evaluation for Automatic Relative Scoring in Machine Translation

Model ReleasesDGX agent

arXiv:2601.18006v2 Announce Type: replace Abstract: We present PEAR (Pairwise Evaluation for Automatic Relative Scoring), a supervised quality estimation (QE) metric family that reframes reference-fre

PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective

Model ReleasesDGX agent

arXiv:2605.28819v1 Announce Type: cross Abstract: Parameter-efficient finetuning (PEFT) has become the standard approach for adapting large language models, yet evaluations largely emphasize downstrea

Personal Visual Memory from Explicit and Implicit Evidence

Model ReleasesDGX agent

arXiv:2605.28806v1 Announce Type: cross Abstract: Long-term memory is increasingly important for personalized AI agents, yet existing benchmarks and methods remain largely text-centric. Even when imag

Personality, Role, and Expressive Style in Large Language Models: An Interactionist Analysis

AgentsDGX agent

arXiv:2605.28037v1 Announce Type: new Abstract: Prompt-based personality control is a key technique for designing large language model (LLM) dialogue agents that behave consistently across social cont

Playing with Words, Improving with Rewards: Training Language Models for Creative Association

ResearchDGX agent

arXiv:2605.27832v1 Announce Type: new Abstract: Large Language Models (LLMs) are being applied to increasingly difficult problems and use cases. To navigate their vast solution spaces effectively, LLM

PrionNER: A Named Entity Recognition Dataset for Prion Disease Biomedical Literature

Model ReleasesDGX agent

arXiv:2605.28375v1 Announce Type: new Abstract: Prion diseases are rare, rapidly progressive, and fatal neurodegenerative disorders that remain difficult to diagnose, particularly in their early stage

Probing Social Identity Bias in Chinese LLMs with Gendered Pronouns and Social Groups

SafetyDGX agent

arXiv:2510.06974v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in user-facing applications, raising concerns that they may reflect and amplify social biases

Prompting Is All You Need: Multi-view Prompting Large Language Models for Aspect-Based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2605.28058v1 Announce Type: new Abstract: Recent work explored the capabilities of Large Language Models (LLMs) in Aspect-Based Sentiment Analysis (ABSA) through few-shot prompting, requiring su

PubMedCausal: A Span-Level Annotated Corpus for Causal Relation Extraction in Biomedical Text

Model ReleasesDGX agent

arXiv:2605.28363v1 Announce Type: new Abstract: Causal relation extraction (CRE) is central to biomedical text mining, but current resources often conflate causal relations with broader associations,

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

SafetyDGX agent

arXiv:2602.15894v2 Announce Type: replace Abstract: In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, exist

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

AgentsDGX agent

arXiv:2605.28003v1 Announce Type: new Abstract: The frontier of mathematics is defined by problems whose solutions are not yet known, yet it remains unclear whether language models can meaningfully en

Rethinking Visual Neglect: Steering via Context-Preference for MLLM Hallucination Mitigation

ResearchDGX agent

arXiv:2605.27993v1 Announce Type: new Abstract: Object hallucination remains a primary obstacle to the reliable deployment of Multimodal Large Language Models (MLLMs). Current inference-time mitigatio

Retrieval, Reward, and Training Protocols: What Matters in Training Search Agents?

SafetyDGX agent

arXiv:2605.27881v1 Announce Type: new Abstract: Search agents powered by large language models can autonomously decompose queries, retrieve information, and synthesize answers through multi-step reaso

ReverseMath: Answer Inversion for Scalable and Verifiable Mathematical Problem Generation

ResearchDGX agent

arXiv:2605.27709v1 Announce Type: new Abstract: Mathematical reasoning benchmarks are vital for evaluating large language models (LLMs), but many are static and repeatedly exposed through public evalu

Risk-aware Selective Prompting for Hallucination Mitigation in Large Vision-Language Models

ResearchDGX agent

arXiv:2605.28123v1 Announce Type: new Abstract: Prompt-based verification is widely used to mitigate hallucinations in large vision-language models (LVLMs), yet when it helps remains poorly understood

RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction

Model ReleasesDGX agent

arXiv:2602.13748v2 Announce Type: replace Abstract: Multimedia Event Extraction (MEE) aims to identify events and their arguments from documents that contain both text and images. It requires groundin

Roles with Rails: Contract-Preserving Role Evolution in Multi-Agent Structured Reasoning

AgentsDGX agent

arXiv:2605.28433v1 Announce Type: new Abstract: Role-based LLM multi-agent systems need adaptive role pools, yet adapting such systems is not merely a matter of prompt optimization: roles often carry

ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains

SafetyDGX agent

arXiv:2605.28014v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) improves the reasoning performance of large language models (LLMs) by providing dense token-level supervision for on-

Self-Consistency via Marginal Sharpening

ResearchDGX agent

arXiv:2605.28142v1 Announce Type: cross Abstract: Inference-time sampling can elicit strong reasoning abilities from language models without additional training. Existing power-sampling methods do so

Self-Improving Language Models with Bidirectional Evolutionary Search

AgentsDGX agent

arXiv:2605.28814v1 Announce Type: new Abstract: Search has been proposed as an effective method for self-improving language models and agentic systems, both for post-training sample generation and for

Sentence Curve Language Models

Local AiDGX agent

arXiv:2602.01807v3 Announce Type: replace Abstract: Language models (LMs) are a central component of modern AI systems, and diffusion language models (DLMs) have recently emerged as a competitive alte

SilentRetrieval: Hijacking Retrieval-Augmented Generation via Semantically-Preserving Adversarial Data Poisoning

ResearchDGX agent

arXiv:2605.28074v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) mitigates LLM hallucinations but introduces a critical vulnerability: corpus integrity. We present SilentRetrieva

Simorgh at SemEval-2026 task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering

Model ReleasesDGX agent

arXiv:2605.27636v1 Announce Type: new Abstract: Although Large Language Models (LLMs) demonstrate excellent capabilities and performance for general reasoning tasks within the general public domain, t

Skill-as-Pseudocode: Refactoring Skill Libraries to Pseudocode for LLM Agents

AgentsDGX agent

arXiv:2605.27955v1 Announce Type: cross Abstract: Markdown skill libraries for LLM agents ship as free-form prose, forcing the agent to re-derive both the input schema and the concrete invocation synt

Skill0.5: Joint Skill Internalization and Utilization for Out-of-Distribution Generalization in Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2605.28424v1 Announce Type: new Abstract: Equipping large language models with explicit skills has emerged as a promising paradigm for enabling autonomous agents to solve complex tasks. Agent sk

Soft-SVeRL: Self-Verified Reinforcement Learning with Soft Rewards

SafetyDGX agent

arXiv:2605.28561v1 Announce Type: new Abstract: Reinforcement Learning from Verifiable Rewards (RLVR) has improved language models in domains such as mathematics and code, where correctness can be che

Stance Detection in Prediction Markets: Addressing Imbalanced Trader Commentary via Counterfactual Augmentation and Market Context

ResearchDGX agent

arXiv:2605.28745v1 Announce Type: new Abstract: Prediction markets such as Polymarket aggregate crowd beliefs into real-time probability estimates, and the comments traders post beneath each market co

Stop Rewarding Hallucinated Steps: Faithfulness-Aware Step-Level Reinforcement Learning for Small Reasoning Models

ResearchDGX agent

arXiv:2602.05897v2 Announce Type: replace Abstract: As large language models become smaller and more efficient, small reasoning models (SRMs) are crucial for enabling chain-of-thought (CoT) reasoning

SuperValid: Capability-Aligned OOD Validation for Generalizable Downstream Scaling

Model ReleasesDGX agent

arXiv:2605.28179v1 Announce Type: new Abstract: Scaling laws guide large language model training by relating compute to cross-entropy loss, and recent work further extends them to predict downstream b

Supervised Semantic Differential for Cross-Cultural Concept Analysis: A Case Study of Human Affect

SafetyDGX agent

arXiv:2605.28225v1 Announce Type: new Abstract: Cross-cultural comparison of psychological meaning requires methods that go beyond word-level translation and examine how semantic dimensions are organi

← Previous
1…5657585960…129
Next →