AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
21 Apr 2026

Adversarial Humanities Benchmark: Results on Stylistic Robustness in Frontier Model Safety

Model ReleasesDGX agent

arXiv:2604.18487v1 Announce Type: new Abstract: The Adversarial Humanities Benchmark (AHB) evaluates whether model safety refusals survive a shift away from familiar harmful prompt forms. Starting fro

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence

AgentsDGX agent

arXiv:2604.18292v1 Announce Type: cross Abstract: Large language models are increasingly expected to serve as general-purpose agents that interact with external, stateful tool environments. The Model

Agents Explore but Agents Ignore: LLMs Lack Environmental Curiosity

AgentsDGX agent

arXiv:2604.17609v1 Announce Type: new Abstract: LLM-based agents are assumed to integrate environmental observations into their reasoning: discovering highly relevant but unexpected information should


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Agree, Disagree, Explain: Decomposing Human Label Variation in NLI through the Lens of Explanations

SafetyDGX agent

arXiv:2510.16458v2 Announce Type: replace Abstract: Natural Language Inference (NLI) datasets often exhibit human label variation. To better understand these variations, explanation-based approaches a

Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs

Model ReleasesDGX agent

arXiv:2601.13099v2 Announce Type: replace Abstract: Arabic is a highly diglossic language where most daily communication occurs in regional dialects rather than Modern Standard Arabic (MSA). Despite t

Align Documents to Questions: Question-Oriented Document Rewriting for Retrieval-Augmented Generation

SafetyDGX agent

arXiv:2604.17325v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) enhances the factuality of Large Language Models (LLMs) by incorporating retrieved documents and/or generated conte

Aligning Backchannel and Dialogue Context Representations via Contrastive LLM Fine-Tuning

SafetyDGX agent

arXiv:2604.16622v1 Announce Type: new Abstract: Backchannels (e.g., `yeah', `mhm', and `right') are short, non-interruptive feedback signals whose lexical form and prosody jointly convey pragmatic mea

Aligning Language Models for Lyric-to-Melody Generation with Rule-Based Musical Constraints

SafetyDGX agent

arXiv:2604.18489v1 Announce Type: cross Abstract: Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically

Aligning Language Models with Real-time Knowledge Editing

Local AiDGX agent

arXiv:2508.01302v3 Announce Type: replace Abstract: Knowledge editing aims to modify outdated knowledge in language models efficiently while retaining their original capabilities. Mainstream datasets

Alignment Data Map for Efficient Preference Data Selection and Diagnosis

SafetyDGX agent

arXiv:2505.23114v3 Announce Type: replace Abstract: Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and ineffic

AlphaContext: An Evolutionary Tree-based Psychometric Context Generator for Creativity Assessment

ApplicationsDGX agent

arXiv:2604.18398v1 Announce Type: new Abstract: Creativity has become a core competence in the era of LLMs and human-AI collaboration, underpinning innovation in real-world problem solving. Crucially,

Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning

Model ReleasesDGX agent

arXiv:2602.11161v2 Announce Type: replace-cross Abstract: The web's information ecosystem demands fact-checking systems that are both scalable and epistemically trustworthy. Automated approaches offer

An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal

ResearchDGX agent

arXiv:2604.18293v1 Announce Type: new Abstract: Surprisal theory hypothesizes that the difficulty of human sentence processing increases linearly with surprisal, the negative log-probability of a word

An Exploration of Mamba for Speech Self-Supervised Models

ResearchDGX agent

arXiv:2506.12606v2 Announce Type: replace Abstract: While Mamba has demonstrated strong performance in language modeling, its potential as a speech self-supervised learning (SSL) model remains underex

AnchorMem: Anchored Facts with Associative Contexts for Building Memory in Large Language Models

Model ReleasesDGX agent

arXiv:2604.17377v1 Announce Type: new Abstract: While large language models have achieved remarkable performance in complex tasks, they still need a memory system to utilize historical experience in l

Annotating Dimensions of Social Perception in Text: A Sentence-Level Dataset of Warmth and Competence

ResearchDGX agent

arXiv:2601.06316v2 Announce Type: replace Abstract: Warmth (W) (often further broken down intoTrust (T) and Sociability (S)) and Competence (C) are central dimensions along which people evaluate indiv

Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning

ResearchDGX agent

arXiv:2604.16332v1 Announce Type: cross Abstract: We find that LoRA fine-tuning exhibits un-learning on contested examples: items with high annotator disagreement show increasing loss during training,

Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems

AgentsDGX agent

arXiv:2604.17487v1 Announce Type: new Abstract: Agentic systems often fail not by being entirely wrong, but by being too precise: a response may be generally useful while particular claims exceed what

ArbGraph: Conflict-Aware Evidence Arbitration for Reliable Long-Form Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2604.18362v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) remains unreliable in long-form settings, where retrieved evidence is noisy or contradictory, making it difficult f

Arch: An AI-Native Hardware Description Language for Register-Transfer Clocked Hardware Design

SafetyDGX agent

arXiv:2604.05983v2 Announce Type: replace-cross Abstract: We present Arch (AI-native Register-transfer Clocked Hardware), a hardware description language for micro-architecture specification and AI-as

Are Emotion and Rhetoric Neurons in LLM? Neuron Recognition and Adaptive Masking for Emotion-Rhetoric Prediction Steering

ResearchDGX agent

arXiv:2604.17255v1 Announce Type: new Abstract: Accurate comprehension and controllable generation of emotion and rhetoric are pivotal for enhancing the reasoning capabilities of large language models

Are they lovers or friends? Evaluating LLMs' Social Reasoning in English and Korean Dialogues

ApplicationsDGX agent

arXiv:2510.19028v3 Announce Type: replace Abstract: As LLMs are increasingly deployed in real-world interactions, their social reasoning in interpersonal communication becomes critical. To explore the

ArgBench: Benchmarking LLMs on Computational Argumentation Tasks

Model ReleasesDGX agent

arXiv:2604.17366v1 Announce Type: new Abstract: Argumentation skills are an essential toolkit for large language models (LLMs). These skills are crucial in various use cases, including self-reflection

Argument Reconstruction as Supervision for Critical Thinking in LLMs

TutorialsDGX agent

arXiv:2603.17432v2 Announce Type: replace Abstract: To think critically about arguments, human learners are trained to identify, reconstruct, and evaluate arguments. Argument reconstruction is especia

ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data

Model ReleasesDGX agent

arXiv:2604.17663v1 Announce Type: cross Abstract: Constitution-conditioned post-training can be analysed as a structured perturbation of a model's learned representational geometry. We introduce ATLAS

Audio-DeepThinker: Progressive Reasoning-Aware Reinforcement Learning for High-Quality Chain-of-Thought Emergence in Audio Language Models

SafetyDGX agent

arXiv:2604.18187v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) have made significant progress in audio understanding, yet they primarily operate as perception-and-answer systems

Auditing Support Strategies in LLMs through Grounded Multi-Turn Social Simulation

Model ReleasesDGX agent

arXiv:2604.17079v1 Announce Type: new Abstract: When users seek social support from chatbots, they disclose their situation gradually, yet most evaluations of supportive LLMs rely on single-turn, full

AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction

SafetyDGX agent

arXiv:2510.15339v3 Announce Type: replace Abstract: Building effective knowledge graphs (KGs) for Retrieval-Augmented Generation (RAG) is pivotal for advancing question answering (QA) systems. However

Automatic Slide Updating with User-Defined Dynamic Templates and Natural Language Instructions

Model ReleasesDGX agent

arXiv:2604.17894v1 Announce Type: new Abstract: Presentation slides are a primary medium for data-driven reporting, yet keeping complex, analytics-style decks up to date remains labor-intensive. Exist

Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan

ApplicationsDGX agent

arXiv:2603.26248v2 Announce Type: replace Abstract: Language endangerment poses a major challenge to linguistic diversity worldwide, and technological advances have opened new avenues for documentatio

AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

ResearchDGX agent

arXiv:2510.14738v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) have rapidly advanced from perception tasks to complex multi-step reasoning, yet reinforcement learning wit

BASIL: Bayesian Assessment of Sycophancy in LLMs

ApplicationsDGX agent

arXiv:2508.16846v5 Announce Type: replace-cross Abstract: Sycophancy (overly agreeable or flattering behavior) poses a fundamental challenge for human-AI collaboration, particularly in high-stakes dec

Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report

ResearchDGX agent

arXiv:2604.17707v1 Announce Type: new Abstract: Clinical personality assessment screens response validity before interpreting substantive scales. LLM evaluation does not. We apply the validity scaling

BEFT: Bias-Efficient Fine-Tuning of Language Models in Low-Data Regimes

Model ReleasesDGX agent

arXiv:2509.15974v2 Announce Type: replace Abstract: Fine-tuning the bias terms of large language models (LLMs) has the potential to achieve unprecedented parameter efficiency while maintaining competi

BenchMarker: An Education-Inspired Toolkit for Highlighting Flaws in Multiple-Choice Benchmarks

Model ReleasesDGX agent

arXiv:2602.06221v2 Announce Type: replace Abstract: Multiple-choice question answering (MCQA) is standard in NLP, but benchmarks lack rigorous quality control. We present BenchMarker, an education-ins

Benchmarking Real-Time Question Answering via Executable Code Workflows

Model ReleasesDGX agent

arXiv:2604.16349v1 Announce Type: cross Abstract: Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are

BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture

Model ReleasesDGX agent

arXiv:2511.03180v2 Announce Type: replace Abstract: As multilingual Large Language Models (LLMs) gain traction across South Asia, their alignment with local ethical norms, particularly for Bengali, sp

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing SubjectiveNLP Tasks

ResearchDGX agent

arXiv:2604.17022v1 Announce Type: new Abstract: Subjective NLP datasets typically aggregate annotator judgments into a single gold label, making it difficult to diagnose whether disagreement reflects

Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation

ResearchDGX agent

arXiv:2604.17574v1 Announce Type: new Abstract: Distractor generation (DG) remains a labor-intensive task that still significantly depends on domain experts. The task focuses on generating plausible y

Beyond 'I Don't Know': Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty

Model ReleasesDGX agent

arXiv:2604.17293v1 Announce Type: new Abstract: Reliable Large Language Models (LLMs) should abstain when confidence is insufficient. However, prior studies often treat refusal as a generic 'I don't k

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization

SafetyDGX agent

arXiv:2604.17188v1 Announce Type: new Abstract: Multi-role dialogue summarization requires modeling complex interactions among multiple speakers while preserving role-specific information and factual

Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection

Model ReleasesDGX agent

arXiv:2604.18248v1 Announce Type: cross Abstract: Current open-source prompt-injection detectors converge on two architectural choices: regular-expression pattern matching and fine-tuned transformer c

Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation

Model ReleasesDGX agent

arXiv:2604.18169v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for creative tasks such as literary translation. Yet translational creativity remains underexplored a

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

ResearchDGX agent

arXiv:2604.17020v1 Announce Type: new Abstract: Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining

ResearchDGX agent

arXiv:2511.21613v2 Announce Type: replace Abstract: Incorporating metadata in Large Language Models (LLMs) pretraining has recently emerged as a promising approach to accelerate training. However prio

Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text

Model ReleasesDGX agent

arXiv:2604.17108v1 Announce Type: new Abstract: Coreference Resolution (CR) is a fundamental NLP task critical for long-form tasks as information extraction, summarization, and many business applicati

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources

ResearchDGX agent

arXiv:2604.18423v1 Announce Type: new Abstract: India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks

BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories

SafetyDGX agent

arXiv:2604.17008v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to generate narrative content, including children's stories, which play an important role in social a

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation

Model ReleasesDGX agent

arXiv:2602.07954v4 Announce Type: replace Abstract: As Large Language Models (LLMs) become increasingly deployed in Polish language applications, the need for efficient and accurate content safety cla

Bolzano: Case Studies in LLM-Assisted Mathematical Research

AgentsDGX agent

arXiv:2604.16989v1 Announce Type: new Abstract: We report new results on six problems in mathematics and theoretical computer science, produced with the assistance of Bolzano, an open-source multi-age

Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction

ResearchDGX agent

arXiv:2604.16370v1 Announce Type: new Abstract: Decoding natural language from non-invasive electroencephalography (EEG) remains fundamentally limited by low signal-to-noise ratio and restricted infor

BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation

SafetyDGX agent

arXiv:2602.23580v2 Announce Type: replace Abstract: In the field of educational assessment, automated scoring systems increasingly rely on deep learning and large language models (LLMs). However, thes

Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languages

Local AiDGX agent

arXiv:2508.14913v4 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant capabilities in solving mathematical problems expressed in natural language. However, mul

Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling

Model ReleasesDGX agent

arXiv:2604.17794v1 Announce Type: new Abstract: The democratization of ubiquitous AI hinges on deploying sophisticated reasoning capabilities on resource-constrained devices. However, Small Language M

Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data

Model ReleasesDGX agent

arXiv:2601.11038v2 Announce Type: replace Abstract: We study the reasoning behavior of large language models (LLMs) under limited computation budgets. In such settings, producing useful partial soluti

Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA

ApplicationsDGX agent

arXiv:2604.17316v1 Announce Type: new Abstract: Safe clinical deployment of Large Language Models (LLMs) requires not only high accuracy but also robust uncertainty calibration to ensure models defer

Calibrating Model-Based Evaluation Metrics for Summarization

ResearchDGX agent

arXiv:2604.17200v1 Announce Type: new Abstract: Recent advances in summary evaluation are based on model-based metrics to assess quality dimensions, such as completeness, conciseness, and faithfulness

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

SafetyDGX agent

arXiv:2510.08986v2 Announce Type: replace Abstract: We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives ann

CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval

Model ReleasesDGX agent

arXiv:2601.17230v2 Announce Type: replace Abstract: Automated Fact-Checking has largely focused on verifying general knowledge against static corpora, overlooking high-stakes domains like law where tr

Cat-DPO: Category-Adaptive Safety Alignment

SafetyDGX agent

arXiv:2604.17299v1 Announce Type: new Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusin

← Previous
1…104105106107108…129
Next →