AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
20 Apr 2026

LLMSniffer: Detecting LLM-Generated Code via GraphCodeBERT and Supervised Contrastive Learning

Model ReleasesDGX agent

arXiv:2604.16058v1 Announce Type: cross Abstract: The rapid proliferation of Large Language Models (LLMs) in software development has made distinguishing AI-generated code from human-written code a cr

Measuring the Semantic Structure and Evolution of Conspiracy Theories

ResearchDGX agent

arXiv:2603.26062v2 Announce Type: replace Abstract: Research on conspiracy theories has largely focused on belief formation, exposure, and diffusion, while paying less attention to how their meanings

MemEvoBench: Benchmarking Memory MisEvolution in LLM Agents

Model ReleasesDGX agent

arXiv:2604.15774v1 Announce Type: new Abstract: Equipping Large Language Models (LLMs) with persistent memory enhances interaction continuity and personalization but introduces new safety risks. Speci


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

MUSCAT: MUltilingual, SCientific ConversATion Benchmark

Model ReleasesDGX agent

arXiv:2604.15929v1 Announce Type: new Abstract: The goal of multilingual speech technology is to facilitate seamless communication between individuals speaking different languages, creating the experi

No Universal Courtesy: A Cross-Linguistic, Multi-Model Study of Politeness Effects on LLMs Using the PLUM Corpus

Model ReleasesDGX agent

arXiv:2604.16275v1 Announce Type: new Abstract: This paper explores the response of Large Language Models (LLMs) to user prompts with different degrees of politeness and impoliteness. The Politeness T

Olmo Hybrid: From Theory to Practice and Back

Model ReleasesDGX agent

arXiv:2604.03444v3 Announce Type: replace-cross Abstract: Recent work has demonstrated the potential of non-transformer language models, especially linear recurrent neural networks (RNNs) and hybrid m

On the Rejection Criterion for Proxy-based Test-time Alignment

SafetyDGX agent

arXiv:2604.16146v1 Announce Type: new Abstract: Recent works proposed test-time alignment methods that rely on a small aligned model as a proxy that guides the generation of a larger base (unaligned)

Optimizing Korean-Centric LLMs via Token Pruning

Model ReleasesDGX agent

arXiv:2604.16235v1 Announce Type: new Abstract: This paper presents a systematic benchmark of state-of-the-art multilingual large language models (LLMs) adapted via token pruning - a compression techn

Predicting Where Steering Vectors Succeed

Model ReleasesDGX agent

arXiv:2604.15557v1 Announce Type: cross Abstract: Steering vectors work for some concepts and layers but fail for others, and practitioners have no way to predict which setting applies before running

Preference Estimation via Opponent Modeling in Multi-Agent Negotiation

Model ReleasesDGX agent

arXiv:2604.15687v1 Announce Type: new Abstract: Automated negotiation in complex, multi-party and multi-issue settings critically depends on accurate opponent modeling. However, conventional numerical

Pruning Unsafe Tickets: A Resource-Efficient Framework for Safer and More Robust LLMs

Model ReleasesDGX agent

arXiv:2604.15780v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in real-world applications, but even aligned models such as Mistral and LLaVA still exhibit unsafe b

Qwen3.5-Omni Technical Report

Model ReleasesDGX agent

arXiv:2604.15804v1 Announce Type: new Abstract: In this work, we present Qwen3.5-Omni, the latest advancement in the Qwen-Omni model family. Representing a significant evolution over its predecessor,

RAGognizer: Hallucination-Aware Fine-Tuning via Detection Head Integration

ResearchDGX agent

arXiv:2604.15945v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) is widely used to augment the input to Large Language Models (LLMs) with external information, such as recent or do

RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models

Model ReleasesDGX agent

arXiv:2601.03699v2 Announce Type: replace Abstract: As large language models (LLMs) become integral to safety-critical applications, ensuring their robustness against adversarial prompts is paramount.

RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees

Model ReleasesDGX agent

arXiv:2604.15736v1 Announce Type: cross Abstract: While Multimodal Large Language Models (MLLMs) excel at generic video understanding, their ability to support specialized, rule-grounded decision-maki

Reward Modeling for Scientific Writing Evaluation

ResearchDGX agent

arXiv:2601.11374v2 Announce Type: replace Abstract: Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage

SCHK-HTC: Sibling Contrastive Learning with Hierarchical Knowledge-Aware Prompt Tuning for Hierarchical Text Classification

Model ReleasesDGX agent

arXiv:2604.15998v1 Announce Type: new Abstract: Few-shot Hierarchical Text Classification (few-shot HTC) is a challenging task that involves mapping texts to a predefined tree-structured label hierarc

Sentiment Analysis of German Sign Language Fairy Tales

ResearchDGX agent

arXiv:2604.16138v1 Announce Type: new Abstract: We present a dataset and a model for sentiment analysis of German sign language (DGS) fairy tales. First, we perform sentiment analysis for three levels

SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space

ResearchDGX agent

arXiv:2504.16315v4 Announce Type: replace-cross Abstract: The complexity of Sign Language (SL) data processing brings many challenges. The current approach to recognition of SL signs aims to translate

SIMMER: Cross-Modal Food Image--Recipe Retrieval via MLLM-Based Embedding

SafetyDGX agent

arXiv:2604.15628v1 Announce Type: cross Abstract: Cross-modal retrieval between food images and recipe texts is an important task with applications in nutritional management, dietary logging, and cook

Skill-RAG: Failure-State-Aware Retrieval Augmentation via Hidden-State Probing and Skill Routing

SafetyDGX agent

arXiv:2604.15771v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) has emerged as a foundational paradigm for grounding large language models in external knowledge. While adaptive re

Stochasticity in Tokenisation Improves Robustness

Model ReleasesDGX agent

arXiv:2604.16037v1 Announce Type: new Abstract: The widespread adoption of large language models (LLMs) has increased concerns about their robustness. Vulnerabilities in perturbations of tokenisation

SwanNLP at SemEval-2026 Task 5: An LLM-based Framework for Plausibility Scoring in Narrative Word Sense Disambiguation

Model ReleasesDGX agent

arXiv:2604.16262v1 Announce Type: new Abstract: Recent advances in language models have substantially improved Natural Language Understanding (NLU). Although widely used benchmarks suggest that Large

Target-Oriented Pretraining Data Selection via Neuron-Activated Graph

ResearchDGX agent

arXiv:2604.15706v1 Announce Type: new Abstract: Everyday tasks come with a target, and pretraining models around this target is what turns them into experts. In this paper, we study target-oriented la

The Metacognitive Monitoring Battery: A Cross-Domain Benchmark for LLM Self-Monitoring

Model ReleasesDGX agent

arXiv:2604.15702v1 Announce Type: new Abstract: We introduce a cross-domain behavioural assay of monitoring-control coupling in LLMs, grounded in the Nelson and Narens (1990) metacognitive framework a

Think Multilingual, Not Harder: A Data-Efficient Framework for Teaching Reasoning Models to Code-Switch

ResearchDGX agent

arXiv:2604.15490v1 Announce Type: new Abstract: Recent developments in reasoning capabilities have enabled large language models to solve increasingly complex mathematical, symbolic, and logical tasks

TRIDENT: Enhancing Large Language Model Safety with Tri-Dimensional Diversified Red-Teaming Data Synthesis

Model ReleasesDGX agent

arXiv:2505.24672v2 Announce Type: replace Abstract: Large Language Models (LLMs) excel in various natural language processing tasks but remain vulnerable to generating harmful content or being exploit

TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models

ResearchDGX agent

arXiv:2604.15756v1 Announce Type: new Abstract: Vision-language models (VLMs) such as CLIP exhibit strong Out-of-distribution (OOD) detection capabilities by aligning visual and textual representation

Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis and Interpretation

ResearchDGX agent

arXiv:2511.02626v3 Announce Type: replace Abstract: Prior works have shown that fine-tuning on new knowledge can induce factual hallucinations in large language models (LLMs), leading to incorrect out

UsefulBench: Towards Decision-Useful Information as a Target for Information Retrieval

SafetyDGX agent

arXiv:2604.15827v1 Announce Type: cross Abstract: Conventional information retrieval is concerned with identifying the relevance of texts for a given query. Yet, the conventional definition of relevan

Whose Facts Win? LLM Source Preferences under Knowledge Conflicts

SafetyDGX agent

arXiv:2601.03746v3 Announce Type: replace Abstract: As large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their beh

WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback

SafetyDGX agent

arXiv:2408.15549v4 Announce Type: replace Abstract: As large language models (LLMs) continue to advance, aligning these models with human preferences has emerged as a critical challenge. Traditional a

Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting

Model ReleasesDGX agent

arXiv:2510.17210v3 Announce Type: replace Abstract: The increase in computing power and the necessity of AI-assisted decision-making boost the growing application of large language models (LLMs). Alon

17 Apr 2026

A Linguistics-Aware LLM Watermarking via Syntactic Predictability

ResearchDGX agent

arXiv:2510.13829v3 Announce Type: replace Abstract: As large language models (LLMs) continue to advance rapidly, reliable governance tools have become critical. Publicly verifiable watermarking is par

AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization

Model ReleasesDGX agent

arXiv:2511.15915v2 Announce Type: replace-cross Abstract: We present AccelOpt, a self-improving large language model (LLM) agentic system that autonomously optimizes kernels for emerging AI acclerator

Acceptance Dynamics Across Cognitive Domains in Speculative Decoding

Model ReleasesDGX agent

arXiv:2604.14682v1 Announce Type: cross Abstract: Speculative decoding accelerates large language model (LLM) inference. It uses a small draft model to propose a tree of future tokens. A larger target

ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints

Model ReleasesDGX agent

arXiv:2604.14902v1 Announce Type: cross Abstract: Intelligent embodied agents should not simply follow instructions, as real-world environments often involve unexpected conditions and exceptions. Howe

Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference

ResearchDGX agent

arXiv:2601.07667v2 Announce Type: replace Abstract: Due to the prevalence of large language models (LLMs), key-value (KV) cache reduction for LLM inference has received remarkable attention. Among num

AdaSplash-2: Faster Differentiable Sparse Attention

HardwareDGX agent

arXiv:2604.15180v1 Announce Type: cross Abstract: Sparse attention has been proposed as a way to alleviate the quadratic cost of transformers, a central bottleneck in long-context training. A promisin

AIM: Asymmetric Information Masking for Visual Question Answering Continual Learning

ResearchDGX agent

arXiv:2604.14779v1 Announce Type: cross Abstract: In continual visual question answering (VQA), existing Continual Learning (CL) methods are mostly built for symmetric, unimodal architectures. However

An Underexplored Frontier: Large Language Models for Rare Disease Patient Education and Communication -- A scoping review

ApplicationsDGX agent

arXiv:2604.14179v1 Announce Type: new Abstract: Rare diseases affect over 300 million people worldwide and are characterized by complex care pathways, limited clinical expertise, and substantial unmet

Anonpsy: A Graph-Based Framework for Structure-Preserving De-identification of Psychiatric Narratives

Model ReleasesDGX agent

arXiv:2601.13503v2 Announce Type: replace Abstract: Psychiatric narratives encode patient identity not only through explicit identifiers but also through idiosyncratic life events embedded in their cl

APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI

AgentsDGX agent

arXiv:2604.14362v1 Announce Type: new Abstract: Large language models still struggle with reliable long-term conversational memory: simply enlarging context windows or applying naive retrieval often i

Attention to Mamba: A Recipe for Cross-Architecture Distillation

TutorialsDGX agent

arXiv:2604.14191v1 Announce Type: new Abstract: State Space Models (SSMs) such as Mamba have become a popular alternative to Transformer models, due to their reduced memory consumption and higher thro

Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models

ResearchDGX agent

arXiv:2508.15396v2 Announce Type: replace Abstract: The increasing adoption of large language models (LLMs) has raised serious concerns about their reliability and trustworthiness. As a result, a grow

Benchmarking Linguistic Adaptation in Comparable-Sized LLMs: A Study of Llama-3.1-8B, Mistral-7B-v0.1, and Qwen3-8B on Romanized Nepali

Model ReleasesDGX agent

arXiv:2604.14171v1 Announce Type: new Abstract: Romanized Nepali, the Nepali language written in the Latin alphabet, is the dominant medium for informal digital communication in Nepal, yet it remains

Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation

AgentsDGX agent

arXiv:2601.07338v2 Announce Type: replace Abstract: Large Language Models (LLMs) have significantly advanced Machine Translation (MT), applying them to linguistically complex domains-such as Social Ne

Beyond Translation: Evaluating Mathematical Reasoning Capabilities of LLMs in Sinhala and Tamil

ResearchDGX agent

arXiv:2602.14517v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved strong results in mathematical reasoning, and are increasingly deployed as tutoring and learning support

BiCon-Gate: Consistency-Gated De-colloquialisation for Dialogue Fact-Checking

Model ReleasesDGX agent

arXiv:2604.14389v1 Announce Type: new Abstract: Automated fact-checking in dialogue involves multi-turn conversations where colloquial language is frequent yet understudied. To address this gap, we pr

Blinded Multi-Rater Comparative Evaluation of a Large Language Model and Clinician-Authored Responses in CGM-Informed Diabetes Counseling

SafetyDGX agent

arXiv:2604.15124v1 Announce Type: new Abstract: Continuous glucose monitoring (CGM) is central to diabetes care, but explaining CGM patterns clearly and empathetically remains time-intensive. Evidence

BoundRL: Efficient Structured Text Segmentation through Reinforced Boundary Generation

SafetyDGX agent

arXiv:2510.20151v2 Announce Type: replace Abstract: Structured texts refer to texts containing structured elements beyond plain texts, such as code snippets and placeholders. Such structured texts inc

CAMO: An Agentic Framework for Automated Causal Discovery from Micro Behaviors to Macro Emergence in LLM Agent Simulations

AgentsDGX agent

arXiv:2604.14691v1 Announce Type: cross Abstract: LLM-empowered agent simulations are increasingly used to study social emergence, yet the micro-to-macro causal mechanisms behind macro outcomes often

Can Large Language Models Detect Methodological Flaws? Evidence from Gesture Recognition for UAV-Based Rescue Operation Based on Deep Learning

ApplicationsDGX agent

arXiv:2604.14161v1 Announce Type: new Abstract: Reliable evaluation is essential in machine learning research, yet methodological flaws-particularly data leakage-continue to undermine the validity of

CausalDetox: Causal Head Selection and Intervention for Language Model Detoxification

Model ReleasesDGX agent

arXiv:2604.14602v1 Announce Type: new Abstract: Large language models (LLMs) frequently generate toxic content, posing significant risks for safe deployment. Current mitigation strategies often degrad

CausalEmbed: Auto-Regressive Multi-Vector Generation in Latent Space for Visual Document Embedding

TutorialsDGX agent

arXiv:2601.21262v3 Announce Type: replace Abstract: Although Multimodal Large Language Models (MLLMs) have shown remarkable potential in Visual Document Retrieval (VDR) through generating high-quality

Challenges in Translating Technical Lectures: Insights from the NPTEL

ResearchDGX agent

arXiv:2602.08698v2 Announce Type: replace Abstract: This study examines the practical applications and methodological implications of Machine Translation in Indian Languages, specifically Bangla, Mala

Chinese Essay Rhetoric Recognition Using LoRA, In-context Learning and Model Ensemble

ApplicationsDGX agent

arXiv:2604.14167v1 Announce Type: new Abstract: Rhetoric recognition is a critical component in automated essay scoring. By identifying rhetorical elements in student writing, AI systems can better as

Chinese Language Is Not More Efficient Than English in Vibe Coding: A Preliminary Study on Token Cost and Problem-Solving Rate

Model ReleasesDGX agent

arXiv:2604.14210v1 Announce Type: new Abstract: A claim has been circulating on social media and practitioner forums that Chinese prompts are more token-efficient than English for LLM coding tasks, po

Chronological Knowledge Retrieval: A Retrieval-Augmented Generation Approach to Construction Project Documentation

ResearchDGX agent

arXiv:2604.14169v1 Announce Type: new Abstract: In large-scale construction projects, the continuous evolution of decisions generates extensive records, most often captured in meeting minutes. Since d

ClimateCause: Complex and Implicit Causal Structures in Climate Reports

SafetyDGX agent

arXiv:2604.14856v1 Announce Type: new Abstract: Understanding climate change requires reasoning over complex causal networks. Yet, existing causal discovery datasets predominantly capture explicit, di

← Previous
1…112113114115116…128
Next →