AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
21 Apr 2026

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs

ResearchDGX agent

arXiv:2602.11528v2 Announce Type: replace-cross Abstract: Recent studies have shown that large language models (LLMs) can infer private user attributes (e.g., age, location, gender) from user-generate

Structure-Aware Diversity Pursuit as an AI Safety Strategy against Homogenization

SafetyDGX agent

arXiv:2601.06116v2 Announce Type: replace-cross Abstract: Generative AI models reproduce the biases in the training data and can further amplify them through mode collapse. We refer to the resulting h

Style over Story: Measuring LLM Narrative Preferences via Structured Selection

ResearchDGX agent

arXiv:2510.02025v4 Announce Type: replace Abstract: We introduce a constraint-selection-based experiment design for measuring narrative preferences of Large Language Models (LLMs). This design offers


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SynopticBench: Evaluating Vision-Language Models on Generating Weather Forecast Discussions of the Future

SafetyDGX agent

arXiv:2604.16451v1 Announce Type: new Abstract: Recent advances in visual-language models (VLMs) have led to significant improvements in a plethora of complex multimodal tasks like image captioning, r

Synthetic Data Generation for Training Diversified Commonsense Reasoning Models

ResearchDGX agent

arXiv:2603.18361v2 Announce Type: replace Abstract: Conversational agents are required to respond to their users not only with high quality (i.e. commonsense bearing) responses, but also considering m

Synthia: Scalable Grounded Persona Generation from Social Media Data

SafetyDGX agent

arXiv:2507.14922v2 Announce Type: replace Abstract: Persona-driven simulations are increasingly used in computational social science, yet their validity critically depends on the fidelity of the under

Systematic Capability Benchmarking of Frontier Large Language Models for Offensive Cyber Tasks

Model ReleasesDGX agent

arXiv:2604.17159v1 Announce Type: cross Abstract: We present, to our knowledge, the most comprehensive cross-model evaluation of LLM agents on offensive cybersecurity tasks, benchmarking 10 frontier m

Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation

ResearchDGX agent

arXiv:2510.09671v2 Announce Type: replace Abstract: Table Question Answering (TQA) aims to answer natural language questions about tabular data, often accompanied by additional contexts such as text p

Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation via MCTS-Guided Reasoning Reconstruction

ResearchDGX agent

arXiv:2508.11184v2 Announce Type: replace Abstract: Distractors-incorrect yet plausible answer choices in multiple-choice questions (MCQs)-are vital in educational assessments, as they help identify s

Task Matters: Knowledge Requirements Shape LLM Responses to Context-Memory Conflict

SafetyDGX agent

arXiv:2506.06485v4 Announce Type: replace Abstract: Large language models (LLMs) draw on both contextual information and parametric memory, yet these sources can conflict. Prior studies have largely e

Test-Time Reasoners Are Strategic Multiple-Choice Test-Takers

ResearchDGX agent

arXiv:2510.07761v2 Announce Type: replace Abstract: Large language models (LLMs) now give reasoning before answering, excelling in tasks like multiple-choice question answering (MCQA). Yet, a concern

TextTIGER: Text-based Intelligent Generation with Entity Prompt Refinement for Text-to-Image Generation

ResearchDGX agent

arXiv:2504.18269v2 Announce Type: replace Abstract: When generating images from prompts that include specific entities, the model must retain as much entity-specific knowledge as possible. However, th

The Cognitive Penalty: Ablating System 1 and System 2 Reasoning in Edge-Native SLMs for Decentralized Consensus

Model ReleasesDGX agent

arXiv:2604.16913v1 Announce Type: cross Abstract: Decentralized Autonomous Organizations (DAOs) are inclined explore Small Language Models (SLMs) as edge-native constitutional firewalls to vet proposa

The Consensus Trap: Rescuing Multi-Agent LLMs from Adversarial Majorities via Token-Level Collaboration

Local AiDGX agent

arXiv:2604.17139v1 Announce Type: new Abstract: Multi-agent large language model (LLM) architectures increasingly rely on response-level aggregation, such as Majority Voting (MAJ), to raise reasoning

The GDN-CC Dataset: Automatic Corpus Clarification for AI-enhanced Democratic Citizen Consultations

ResearchDGX agent

arXiv:2601.14944v3 Announce Type: replace Abstract: LLMs are ubiquitous in modern NLP, and while their applicability extends to texts produced for democratic activities such as online deliberations or

The Geometric Canary: Predicting Steerability and Detecting Drift via Representational Stability

Model ReleasesDGX agent

arXiv:2604.17698v1 Announce Type: cross Abstract: Reliable deployment of language models requires two capabilities that appear distinct but share a common geometric foundation: predicting whether a mo

The Illusion of Insight in Reasoning Models

Model ReleasesDGX agent

arXiv:2601.00514v2 Announce Type: replace-cross Abstract: Do reasoning models have 'Aha!' moments? Prior work suggests that models like DeepSeek-R1-Zero undergo sudden mid-trace realizations that lead

The impact of postediting on AI generative translation in Yemeni context: Translating literary prose by ChatGPT

ResearchDGX agent

arXiv:2604.16704v1 Announce Type: new Abstract: This study examines the role of artificial intelligence in translation, focusing on ChatGPT, specifically ChatGPT-4, and the extent to which human poste

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey

AgentsDGX agent

arXiv:2509.02547v5 Announce Type: replace-cross Abstract: The emergence of agentic reinforcement learning (Agentic RL) marks a paradigm shift from conventional reinforcement learning applied to large

The MediaSpin Dataset: Post-Publication News Headline Edits Annotated for Media Bias

Model ReleasesDGX agent

arXiv:2412.02271v4 Announce Type: replace Abstract: We present MediaSpin, a large-scale language resource capturing how major news outlets modify headlines after publication, and MediaSpin-in-the-Wild

The Provenance Gap in Clinical AI: Evidence-Traceable Temporal Knowledge Graphs for Rare Disease Reasoning

SafetyDGX agent

arXiv:2604.17114v1 Announce Type: new Abstract: Frontier large language models generate clinically accurate outputs, but their citations are often fabricated. We term this the Provenance Gap. We teste

The Role of Vocabularies in Learning Sparse Representations for Ranking

ApplicationsDGX agent

arXiv:2509.16621v2 Announce Type: replace-cross Abstract: Learned Sparse Retrieval (LSR) such as SPLADE has growing interest for effective semantic 1st stage matching while enjoying the efficiency of

The Thin Line Between Comprehension and Persuasion in LLMs

AgentsDGX agent

arXiv:2507.01936v3 Announce Type: replace Abstract: Large language models (LLMs) are excellent at maintaining high-level, convincing dialogue, but it remains unclear whether their persuasive success r

ThinkBrake: Efficient Reasoning via Log-Probability Margin Guided Decoding

ResearchDGX agent

arXiv:2510.00546v5 Announce Type: replace Abstract: Large Reasoning Models (LRMs) allocate substantial inference-time compute to Chain-of-Thought (CoT) reasoning, improving performance on mathematics,

ThreadSumm: Summarization of Nested Discourse Threads Using Tree of Thoughts

ResearchDGX agent

arXiv:2604.17648v1 Announce Type: new Abstract: Summarizing deeply nested discussion threads requires handling interleaved replies, quotes, and overlapping topics, which standard LLM summarizers strug

TLoRA: Task-aware Low Rank Adaptation of Large Language Models

Model ReleasesDGX agent

arXiv:2604.18124v1 Announce Type: new Abstract: Low-Rank Adaptation (LoRA) has become a widely adopted parameter-efficient fine-tuning method for large language models, with its effectiveness largely

TMD-TTS: A Unified Tibetan Multi-Dialect Text-to-Speech Framework for U-Tsang, Amdo and Kham Speech Dataset Generation

ResearchDGX agent

arXiv:2509.18060v2 Announce Type: replace Abstract: Tibetan is a low-resource language with limited parallel speech corpora spanning its three major dialects (U-Tsang, Amdo, and Kham), limiting progre

ToMMeR -- Efficient Entity Mention Detection from Large Language Models

ResearchDGX agent

arXiv:2510.19410v2 Announce Type: replace Abstract: Identifying which text spans refer to entities - mention detection - is both foundational for information extraction and a known performance bottlen

Tool Learning Needs Nothing More Than a Free 8B Language Model

ResearchDGX agent

arXiv:2604.17739v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a prevalent paradigm for training tool calling agents, which typically requires online interactive environments

Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement

SafetyDGX agent

arXiv:2604.06155v2 Announce Type: replace-cross Abstract: Whether Large Language Models (LLMs) develop coherent internal world models remains a core debate. While conventional Next-Token Prediction (N

Toward Reusability of AI Models Using Dynamic Updates of AI Documentation

SafetyDGX agent

arXiv:2604.17626v1 Announce Type: cross Abstract: This work addresses the challenge of disseminating reusable artificial intelligence (AI) models accompanied by AI documentation (a.k.a., AI model card

Towards Intelligent Legal Document Analysis: CNN-Driven Classification of Case Law Texts

ApplicationsDGX agent

arXiv:2604.17674v1 Announce Type: new Abstract: Legal practitioners and judicial institutions face an ever-growing volume of case-law documents characterised by formalised language, lengthy sentence s

Towards Self-Improving Error Diagnosis in Multi-Agent Systems

AgentsDGX agent

arXiv:2604.17658v1 Announce Type: cross Abstract: Large Language Model (LLM)-based Multi-Agent Systems (MAS) enable complex problem-solving but introduce significant debugging challenges, characterize

ToxiFrench: Benchmarking and Enhancing Language Models via CoT Fine-Tuning for French Toxicity Detection

Model ReleasesDGX agent

arXiv:2508.11281v3 Announce Type: replace Abstract: Detecting toxic content using language models is crucial yet challenging. While substantial progress has been made in English, toxicity detection in

Training for Compositional Sensitivity Reduces Dense Retrieval Generalization

Model ReleasesDGX agent

arXiv:2604.16351v1 Announce Type: cross Abstract: Dense retrieval compresses texts into single embeddings ranked by cosine similarity. While efficient for recall, this interface is brittle for identit

Training Language Models to Use Prolog as a Tool

SafetyDGX agent

arXiv:2512.07407v2 Announce Type: replace Abstract: Language models frequently produce plausible yet incorrect reasoning traces that are difficult to verify. We investigate fine-tuning models to use P

Transition-Matrix Regularization for Next Dialogue Act Prediction in Counselling Conversations

SafetyDGX agent

arXiv:2604.18539v1 Announce Type: new Abstract: This paper studies how empirical dialogue-flow statistics can be incorporated into Next Dialogue Act Prediction (NDAP). A KL regularization term is prop

TriangleMix: Accelerating Prefilling via Decoding-time Contribution Sparsity

Model ReleasesDGX agent

arXiv:2507.21526v3 Announce Type: replace Abstract: Large Language Models (LLMs) incur quadratic attention complexity with input length, creating a major time bottleneck in the prefilling stage. Exist

Triples and Knowledge-Infused Embeddings for Clustering and Classification of Scientific Documents

Model ReleasesDGX agent

arXiv:2601.08841v2 Announce Type: replace Abstract: The increasing volume and complexity of scientific literature demand robust methods for organizing and understanding research documents. In this stu

TSVer: A Benchmark for Fact Verification Against Time-Series Evidence

Model ReleasesDGX agent

arXiv:2511.01101v2 Announce Type: replace Abstract: Reasoning over temporal and numerical data, such as time series, is a crucial aspect of fact-checking. While many systems have recently been develop

TWGuard: A Case Study of LLM Safety Guardrails for Localized Linguistic Contexts

Local AiDGX agent

arXiv:2604.16542v1 Announce Type: cross Abstract: Safety guardrails have become an active area of research in AI safety, aimed at ensuring the appropriate behavior of large language models (LLMs). How

Understanding the Prompt Sensitivity

ResearchDGX agent

arXiv:2604.18389v1 Announce Type: new Abstract: Prompt sensitivity, which refers to how strongly the output of a large language model (LLM) depends on the exact wording of its input prompt, raises con

Unleashing Spatial Reasoning in Multimodal Large Language Models via Textual Representation Guided Reasoning

Model ReleasesDGX agent

arXiv:2603.23404v2 Announce Type: replace-cross Abstract: Existing Multimodal Large Language Models (MLLMs) struggle with 3D spatial reasoning, as they fail to construct structured abstractions of the

User-Assistant Bias in LLMs

Model ReleasesDGX agent

arXiv:2508.15815v3 Announce Type: replace Abstract: Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicit

Using Perspectival Words Is Harder Than Vocabulary Words for Humans and Even More So for Multimodal Language Models

ResearchDGX agent

arXiv:2506.00065v2 Announce Type: replace Abstract: Multimodal language models (MLMs) increasingly demonstrate human-like communication, yet their use of everyday perspectival words remains poorly und

VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis

ResearchDGX agent

arXiv:2509.16538v3 Announce Type: replace-cross Abstract: We propose VC-Inspector, a lightweight, open-source large multimodal model (LMM) for reference-free evaluation of video captions, with a focus

VCORE: Variance-Controlled Optimization-based Reweighting for Chain-of-Thought Supervision

Model ReleasesDGX agent

arXiv:2510.27462v2 Announce Type: replace Abstract: Supervised fine-tuning (SFT) on long chain-of-thought (CoT) trajectories has emerged as a crucial technique for enhancing the reasoning abilities of

VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech

SafetyDGX agent

arXiv:2604.17248v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) are increasingly integrated into daily applications, yet their generative biases remain underexplored. Existing sp

Video-Robin: Autoregressive Diffusion Planning for Intent-Grounded Video-to-Music Generation

Local AiDGX agent

arXiv:2604.17656v1 Announce Type: cross Abstract: Video-to-music (V2M) is the fundamental task of creating background music for an input video. Recent V2M models achieve audiovisual alignment by typic

Vision-Braille: A Curriculum Learning Toolkit and Braille-Chinese Corpus for Braille Translation

ApplicationsDGX agent

arXiv:2407.06048v2 Announce Type: replace Abstract: We present Vision-Braille, the first publicly available end-to-end system for translating Chinese Braille extracted from images into written Chinese

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Model ReleasesDGX agent

arXiv:2505.20279v4 Announce Type: replace-cross Abstract: The rapid advancement of Large Multimodal Models (LMMs) for 2D images and videos has motivated extending these models to understand 3D scenes,

Vocab Diet: Reshaping the Vocabulary of LLMs via Vector Arithmetic

ResearchDGX agent

arXiv:2510.17001v2 Announce Type: replace Abstract: Large language models (LLMs) often encode word-form variation (e.g., walk vs. walked) as linear directions in the embedding space. However, standard

VocabTailor: Dynamic Vocabulary Selection for Downstream Tasks in Small Language Models

Local AiDGX agent

arXiv:2508.15229v3 Announce Type: replace Abstract: Small Language Models (SLMs) provide computational advantages in resource-constrained environments, yet memory limitations remain a critical bottlen

Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception

SafetyDGX agent

arXiv:2604.17475v1 Announce Type: cross Abstract: Small Vision-Language Models (SVLMs) are efficient task controllers but often suffer from visual brittleness and poor tool orchestration. They typical

WeatherArchive-Bench: Benchmarking Retrieval-Augmented Reasoning for Historical Weather Archives

Model ReleasesDGX agent

arXiv:2510.05336v2 Announce Type: replace Abstract: Historical archives on weather events are collections of enduring primary source records that offer rich, untapped narratives of how societies have

What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations

AgentsDGX agent

arXiv:2510.17795v3 Announce Type: replace Abstract: Replicating AI research is a crucial yet challenging task for large language model (LLM) agents. Existing approaches often struggle to generate exec

What makes an entity salient in discourse?

ResearchDGX agent

arXiv:2508.16464v2 Announce Type: replace Abstract: Entities in discourse vary in salience: main participants, objects and locations stay prominent, while others are quickly forgotten, raising questio

When Choices Become Risks: Safety Failures of Large Language Models under Multiple-Choice Constraints

SafetyDGX agent

arXiv:2604.16916v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is primarily evaluated under open-ended generation, where models can mitigate risk by refusing to respo

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

Model ReleasesDGX agent

arXiv:2601.04043v2 Announce Type: replace Abstract: As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to

When Informal Text Breaks NLI: Tokenization Failure, Distribution Shift, and Targeted Mitigations

ResearchDGX agent

arXiv:2604.16787v1 Announce Type: new Abstract: We study how informal surface forms degrade NLI accuracy in ELECTRA-small (14M) and RoBERTa-large (355M) across four transforms applied to SNLI and Mult

← Previous
1…111112113114115…129
Next →