AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
24 Apr 2026

Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions

SafetyDGX agent

arXiv:2604.21871v1 Announce Type: new Abstract: Human moral judgment is context-dependent and modulated by interpersonal relationships. As large language models (LLMs) increasingly function as decisio

Machine learning and digital pragmatics: Which word category influences emoji use most?

ResearchDGX agent

arXiv:2604.21108v1 Announce Type: new Abstract: This study investigates Machine Learning (ML) in the prediction of emojis in Arabic tweets employing the (state-of-the-art) MARBERT model. A corpus of 1

Mapping the Political Discourse in the Brazilian Chamber of Deputies: A Multi-Faceted Computational Approach

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.21897v1 Announce Type: new Abstract: Analyses of legislative behavior often rely on voting records, overlooking the rich semantic and rhetorical content of political speech. In this paper,

MathDuels: Evaluating LLMs as Problem Posers and Solvers

Model ReleasesDGX agent

arXiv:2604.21916v1 Announce Type: new Abstract: As frontier language models attain near-ceiling performance on static mathematical benchmarks, existing evaluations are increasingly unable to different

Measuring Opinion Bias and Sycophancy via LLM-based Coercion

Model ReleasesDGX agent

arXiv:2604.21564v1 Announce Type: new Abstract: Large language models increasingly shape the information people consume: they are embedded in search, consulted for professional advice, deployed as age

Misinformation Span Detection in Videos via Audio Transcripts

ResearchDGX agent

arXiv:2604.21767v1 Announce Type: new Abstract: Online misinformation is one of the most challenging issues lately, yielding severe consequences, including political polarization, attacks on democracy

MKJ at SemEval-2026 Task 9: A Comparative Study of Generalist, Specialist, and Ensemble Strategies for Multilingual Polarization

ResearchDGX agent

arXiv:2604.21370v1 Announce Type: new Abstract: We present a systematic study of multilingual polarization detection across 22 languages for SemEval-2026 Task 9 (Subtask 1), contrasting multilingual g

Multilingual and Domain-Agnostic Tip-of-the-Tongue Query Generation for Simulated Evaluation

Model ReleasesDGX agent

arXiv:2604.21096v1 Announce Type: cross Abstract: Tip-of-the-Tongue (ToT) retrieval benchmarks have largely focused on English, limiting their applicability to multilingual information access. In this

Multilinguality at the Edge: Developing Language Models for the Global South

ResearchDGX agent

arXiv:2604.21637v1 Announce Type: new Abstract: Where and how language models (LMs) are deployed determines who can benefit from them. However, there are several challenges that prevent effective depl

Optimal Aggregation of LLM and PRM Signals for Efficient Test-Time Scaling

ResearchDGX agent

arXiv:2510.13918v2 Announce Type: replace Abstract: Process reward models (PRMs) are a cornerstone of test-time scaling (TTS), designed to verify and select the best responses from large language mode

OptiVerse: A Comprehensive Benchmark towards Optimization Problem Solving

Model ReleasesDGX agent

arXiv:2604.21510v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate remarkable reasoning, complex optimization tasks remain challenging, requiring domain knowledge and robus

Participation and Representation in Local Government Speech

SafetyDGX agent

arXiv:2604.21202v1 Announce Type: cross Abstract: Local government meetings are the most common formal channel through which residents speak directly with elected officials, contest policies, and shap

Phonological Subspace Collapse Is Aetiology-Specific and Cross-Lingually Stable: Evidence from 3,374 Speakers

ResearchDGX agent

arXiv:2604.21706v1 Announce Type: new Abstract: We previously introduced a training-free method for dysarthria severity assessment based on d-prime separability of phonological feature subspaces in fr

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages

ResearchDGX agent

arXiv:2604.21481v1 Announce Type: new Abstract: Crowdsourced pairwise evaluation has emerged as a scalable approach for assessing foundation models. However, applying it to Text to Speech(TTS) introdu

Prefix Parsing is Just Parsing

ResearchDGX agent

arXiv:2604.21191v1 Announce Type: new Abstract: Prefix parsing asks whether an input prefix can be extended to a complete string generated by a given grammar. In the weighted setting, it also provides

ReFACT: A Benchmark for Scientific Confabulation Detection with Positional Error Annotations

Model ReleasesDGX agent

arXiv:2509.25868v3 Announce Type: replace Abstract: The mechanisms underlying scientific confabulation in Large Language Models (LLMs) remain poorly understood. We introduce ReFACT (Reddit False And C

Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms

ResearchDGX agent

arXiv:2604.21882v1 Announce Type: new Abstract: Understanding what kinds of factual knowledge large language models (LLMs) memorize is essential for evaluating their reliability and limitations. Entit

RewardBench 2: Advancing Reward Model Evaluation

Model ReleasesDGX agent

arXiv:2506.01937v2 Announce Type: replace Abstract: Reward models are used throughout the post-training of language models to capture nuanced signals from preference data and provide a training target

Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models

Model ReleasesDGX agent

arXiv:2604.21523v1 Announce Type: cross Abstract: Large Vision-Language Models (VLMs) are increasingly used to evaluate outputs of other models, for image-to-text (I2T) tasks such as visual question a

SemEval-2026 Task 4: Narrative Story Similarity and Narrative Representation Learning

ResearchDGX agent

arXiv:2604.21782v1 Announce Type: new Abstract: We present the shared task on narrative similarity and narrative representation learning - NSNRL (pronounced 'nass-na-rel'). The task operationalizes na

SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding

Model ReleasesDGX agent

arXiv:2510.26615v3 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) extends large language models (LLMs) with external knowledge, but it must balance limited effective context, re

Slot Machines: How LLMs Keep Track of Multiple Entities

ResearchDGX agent

arXiv:2604.21139v1 Announce Type: new Abstract: Language models must bind entities to the attributes they possess and maintain several such binding relationships within a context. We study how multipl

SocraticKG: Knowledge Graph Construction via QA-Driven Fact Extraction

Model ReleasesDGX agent

arXiv:2601.10003v2 Announce Type: replace Abstract: Constructing Knowledge Graphs (KGs) from unstructured text provides a structured framework for knowledge representation and reasoning, yet current L

StegoStylo: Squelching Stylometric Scrutiny through Steganographic Stitching

ResearchDGX agent

arXiv:2601.09056v3 Announce Type: replace-cross Abstract: Stylometry--the identification of an author through analysis of a text's style (i.e., authorship attribution)--serves many constructive purpos

Sub-Token Routing in LoRA for Adaptation and Query-Aware KV Compression

ResearchDGX agent

arXiv:2604.21335v1 Announce Type: cross Abstract: Sub-token routing offers a finer control axis for transformer efficiency than the coarse units used in most prior work, such as tokens, pages, heads,

Subject-level Inference for Realistic Text Anonymization Evaluation

Model ReleasesDGX agent

arXiv:2604.21211v1 Announce Type: new Abstract: Current text anonymization evaluation relies on span-based metrics that fail to capture what an adversary could actually infer, and assumes a single dat

TabSHAP

Local AiDGX agent

arXiv:2604.21120v1 Announce Type: cross Abstract: Large Language Models (LLMs) fine-tuned on serialized tabular data are emerging as powerful alternatives to traditional tree-based models, particularl

The Root Theorem of Context Engineering

Model ReleasesDGX agent

arXiv:2604.20874v1 Announce Type: cross Abstract: Every system that maintains a large language model conversation beyond a single session faces two inescapable constraints: the context window is finit

'This Wasn't Made for Me': Recentering User Experience and Emotional Impact in the Evaluation of ASR Bias

SafetyDGX agent

arXiv:2604.21148v1 Announce Type: new Abstract: Studies on bias in Automatic Speech Recognition (ASR) tend to focus on reporting error rates for speakers of underrepresented dialects, yet less researc

TRACES: Tagging Reasoning Steps for Adaptive Cost-Efficient Early-Stopping

ResearchDGX agent

arXiv:2604.21057v1 Announce Type: new Abstract: The field of Language Reasoning Models (LRMs) has been very active over the past few years with advances in training and inference techniques enabling L

Tuning for TraceTarnish: Techniques, Trends, and Testing Tangible Traits

ResearchDGX agent

arXiv:2512.03465v3 Announce Type: replace-cross Abstract: In this study, we more rigorously evaluated our attack script extit{TraceTarnish}, which leverages adversarial stylometry principles to anonym

UKP_Psycontrol at SemEval-2026 Task 2: Modeling Valence and Arousal Dynamics from Text

ResearchDGX agent

arXiv:2604.21534v1 Announce Type: new Abstract: This paper presents our system developed for SemEval-2026 Task 2. The task requires modeling both current affect and short-term affective change in chro

Unlocking the Power of Large Language Models for Multi-table Entity Matching

ResearchDGX agent

arXiv:2604.21238v1 Announce Type: new Abstract: Multi-table entity matching (MEM) addresses the limitations of dual-table approaches by enabling simultaneous identification of equivalent entities acro

Unveiling Unicode's Unseen Underpinnings in Undermining Authorship Attribution

AgentsDGX agent

arXiv:2508.15840v5 Announce Type: replace-cross Abstract: When using a public communication channel--whether formal or informal, such as commenting or posting on social media--end users have no expect

Weighting What Matters: Boosting Sample Efficiency in Medical Report Generation via Token Reweighting

ResearchDGX agent

arXiv:2604.21082v1 Announce Type: new Abstract: Training vision-language models (VLMs) for medical report generation is often hindered by the scarcity of high-quality annotated data. This work evaluat

When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors

Model ReleasesDGX agent

arXiv:2604.21255v1 Announce Type: new Abstract: Model distillation is a primary driver behind the rapid progress of LLM agents, yet it often leads to behavioral homogenization. Many emerging agents sh

When Bigger Isn't Better: A Comprehensive Fairness Evaluation of Political Bias in Multi-News Summarisation

SafetyDGX agent

arXiv:2604.21309v1 Announce Type: new Abstract: Multi-document news summarisation systems are increasingly adopted for their convenience in processing vast daily news content, making fairness across d

XtraGPT: Context-Aware and Controllable Academic Paper Revision via Human-AI Collaboration

SafetyDGX agent

arXiv:2505.11336v4 Announce Type: replace Abstract: Despite the growing adoption of large language models (LLMs) in academic workflows, their capabilities remain limited in supporting high-quality sci

23 Apr 2026

AFMRL: Attribute-Enhanced Fine-Grained Multi-Modal Representation Learning in E-commerce

ResearchDGX agent

arXiv:2604.20135v1 Announce Type: new Abstract: Multimodal representation is crucial for E-commerce tasks such as identical product retrieval. Large representation models (e.g., VLM2Vec) demonstrate s

Aligning Human-AI-Interaction Trust for Mental Health Support: Survey and Position for Multi-Stakeholders

SafetyDGX agent

arXiv:2604.20166v1 Announce Type: new Abstract: Building trustworthy AI systems for mental health support is a shared priority across stakeholders from multiple disciplines. However, 'trustworthy' rem

Aligning Stuttered-Speech Research with End-User Needs: Scoping Review, Survey, and Guidelines

ResearchDGX agent

arXiv:2604.20535v1 Announce Type: new Abstract: Atypical speech is receiving greater attention in speech technology research, but much of this work unfolds with limited interdisciplinary dialogue. For

All Languages Matter: Understanding and Mitigating Language Bias in Multilingual RAG

SafetyDGX agent

arXiv:2604.20199v1 Announce Type: new Abstract: Multilingual Retrieval-Augmented Generation (mRAG) leverages cross-lingual evidence to ground Large Language Models (LLMs) in global knowledge. However,

Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2604.19974v1 Announce Type: cross Abstract: Large language models can be uncertain yet correct, or confident yet wrong, raising the question of whether their output-level uncertainty and their a

Ask Only When Needed: Proactive Retrieval from Memory and Skills for Experience-Driven Lifelong Agents

SafetyDGX agent

arXiv:2604.20572v1 Announce Type: new Abstract: Online lifelong learning enables agents to accumulate experience across interactions and continually improve on long-horizon tasks. However, existing me

Avoiding Overthinking and Underthinking: Curriculum-Aware Budget Scheduling for LLMs

SafetyDGX agent

arXiv:2604.19780v1 Announce Type: new Abstract: Scaling test-time compute via extended reasoning has become a key paradigm for improving the capabilities of large language models (LLMs). However, exis

Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations

ResearchDGX agent

arXiv:2509.25844v3 Announce Type: replace Abstract: When people query Vision-Language Models (VLMs) but cannot see the accompanying visual context (e.g. for blind and low-vision users), augmenting VLM

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning

Model ReleasesDGX agent

arXiv:2512.15146v3 Announce Type: replace Abstract: Test-time reinforcement learning mitigates the reliance on annotated data by using majority voting results as pseudo-labels, emerging as a complemen

Beyond the Crowd: LLM-Augmented Community Notes for Governing Health Misinformation

Model ReleasesDGX agent

arXiv:2510.11423v3 Announce Type: replace-cross Abstract: Community Notes, the crowd-sourced misinformation governance system on X (formerly Twitter), allows users to flag misleading posts, attach con

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text

Model ReleasesDGX agent

arXiv:2604.20051v1 Announce Type: new Abstract: Self-play has recently emerged as a promising paradigm to train Large Language Models (LLMs). In self-play, the target LLM creates the task input (e.g.,

Breaking the Assistant Mold: Modeling Behavioral Variation in LLM Based Procedural Character Generation

SafetyDGX agent

arXiv:2601.03396v2 Announce Type: replace Abstract: Procedural content generation has enabled vast virtual worlds through levels, maps, and quests, but large-scale character generation remains underex

Chasing the Public Score: User Pressure and Evaluation Exploitation in Coding Agent Workflows

Model ReleasesDGX agent

arXiv:2604.20200v1 Announce Type: new Abstract: Frontier coding agents are increasingly used in workflows where users supervise progress primarily through repeated improvement of a public score, namel

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values

Model ReleasesDGX agent

arXiv:2509.03740v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) like CLIP have shown impressive zero-shot and few-shot learning capabilities across diverse applications. Howeve

Commonsense Knowledge with Negation: A Resource to Enhance Negation Understanding

ResearchDGX agent

arXiv:2604.19921v1 Announce Type: new Abstract: Negation is a common and important semantic feature in natural language, yet Large Language Models (LLMs) struggle when negation is involved in natural

Composition-RL: Compose Your Verifiable Prompts for Reinforcement Learning of Large Language Models

ResearchDGX agent

arXiv:2602.12036v2 Announce Type: replace Abstract: Large-scale verifiable prompts underpin the success of Reinforcement Learning with Verifiable Rewards (RLVR), but they contain many uninformative ex

Construction of a Battery Research Knowledge Graph using a Global Open Catalog

ResearchDGX agent

arXiv:2604.20241v1 Announce Type: new Abstract: Battery research is a rapidly growing and highly interdisciplinary field, making it increasingly difficult to track relevant expertise and identify pote

Continuous Semantic Caching for Low-Cost LLM Serving

ApplicationsDGX agent

arXiv:2604.20021v1 Announce Type: cross Abstract: As Large Language Models (LLMs) become increasingly popular, caching responses so that they can be reused by users with semantically similar queries h

Cooperative Profiles Predict Multi-Agent LLM Team Performance in AI for Science Workflows

Model ReleasesDGX agent

arXiv:2604.20658v1 Announce Type: new Abstract: Multi-agent systems built from teams of large language models (LLMs) are increasingly deployed for collaborative scientific reasoning and problem-solvin

CRAFT: Training-Free Cascaded Retrieval for Tabular QA

Model ReleasesDGX agent

arXiv:2505.14984v2 Announce Type: replace Abstract: Open-Domain Table Question Answering (TQA) involves retrieving relevant tables from a large corpus to answer natural language queries. Traditional d

Decoding Text Spans for Efficient and Accurate Named-Entity Recognition

Local AiDGX agent

arXiv:2604.20447v1 Announce Type: new Abstract: Named Entity Recognition (NER) is a key component in industrial information extraction pipelines, where systems must satisfy strict latency and throughp

Development and Preliminary Evaluation of a Domain-Specific Large Language Model for Tuberculosis Care in South Africa

Model ReleasesDGX agent

arXiv:2604.19776v1 Announce Type: new Abstract: Tuberculosis (TB) is one of the world's deadliest infectious diseases, and in South Africa, it contributes a significant burden to the country's health

← Previous
1…100101102103104…129
Next →