AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
13 Apr 2026

MT-OSC: Path for LLMs that Get Lost in Multi-Turn Conversation

AgentsDGX agent

arXiv:2604.08782v1 Announce Type: new Abstract: Large language models (LLMs) suffer significant performance degradation when user instructions and context are distributed over multiple conversational

Multi-User Large Language Model Agents

AgentsDGX agent

arXiv:2604.08567v1 Announce Type: new Abstract: Large language models (LLMs) and LLM-based agents are increasingly deployed as assistants in planning and decision making, yet most existing systems are

NCL-BU at SemEval-2026 Task 3: Fine-tuning XLM-RoBERTa for Multilingual Dimensional Sentiment Regression

Model ReleasesDGX agent

arXiv:2604.08923v1 Announce Type: new Abstract: Dimensional Aspect-Based Sentiment Analysis (DimABSA) extends traditional ABSA from categorical polarity labels to continuous valence-arousal (VA) regre


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

No Single Best Model for Diversity: Learning a Router for Sample Diversity

ResearchDGX agent

arXiv:2604.02319v2 Announce Type: replace Abstract: When posed with prompts that permit a large number of valid answers, comprehensively generating them is the first step towards satisfying a wide ran

Offline-First LLM Architecture for Adaptive Learning in Low-Connectivity Environments

Local AiDGX agent

arXiv:2603.03339v5 Announce Type: replace-cross Abstract: Artificial intelligence (AI) and large language models (LLMs) are transforming educational technology by enabling conversational tutoring, per

Optimal Multi-bit Generative Watermarking Schemes Under Worst-Case False-Alarm Constraints

ResearchDGX agent

arXiv:2604.08759v1 Announce Type: cross Abstract: This paper considers the problem of multi-bit generative watermarking for large language models under a worst-case false-alarm constraint. Prior work

p1: Better Prompt Optimization with Fewer Prompts

ResearchDGX agent

arXiv:2604.08801v1 Announce Type: cross Abstract: Prompt optimization improves language models without updating their weights by searching for a better system prompt, but its effectiveness varies wide

PaceLLM: Brain-Inspired Large Language Models for Long-Context Understanding

ResearchDGX agent

arXiv:2506.17310v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) demonstrate strong performance across domains, their long-context capabilities are limited by transient neu

PRAGMA: Revolut Foundation Model

ResearchDGX agent

arXiv:2604.08649v1 Announce Type: cross Abstract: Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper presents PRAGMA,

Prototype-Regularized Federated Learning for Cross-Domain Aspect Sentiment Triplet Extraction

ResearchDGX agent

arXiv:2604.09123v1 Announce Type: new Abstract: Aspect Sentiment Triplet Extraction (ASTE) aims to extract all sentiment triplets of aspect terms, opinion terms, and sentiment polarities from a senten

Quantisation Reshapes the Metacognitive Geometry of Language Models

Model ReleasesDGX agent

arXiv:2604.08976v1 Announce Type: new Abstract: We report that model quantisation restructures domain-level metacognitive efficiency in LLMs rather than degrading it uniformly. Evaluating Llama-3-8B-I

Revisiting Anisotropy in Language Transformers: The Geometry of Learning Dynamics

ResearchDGX agent

arXiv:2604.08764v1 Announce Type: new Abstract: Since their introduction, Transformer architectures have dominated Natural Language Processing (NLP). However, recent research has highlighted an inhere

ScheMatiQ: From Research Question to Structured Data through Interactive Schema Discovery

ApplicationsDGX agent

arXiv:2604.09237v1 Announce Type: new Abstract: Many disciplines pose natural-language research questions over large document collections whose answers typically require structured evidence, tradition

Sentiment Classification of Gaza War Headlines: A Comparative Analysis of Large Language Models and Arabic Fine-Tuned BERT Models

Model ReleasesDGX agent

arXiv:2604.08566v1 Announce Type: new Abstract: This study examines how different artificial intelligence architectures interpret sentiment in conflict-related media discourse, using the 2023 Gaza War

SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding

Model ReleasesDGX agent

arXiv:2507.20185v2 Announce Type: replace Abstract: Session history is a common way of recording user interacting behaviors throughout a browsing activity with multiple products. For example, if an us

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos

Model ReleasesDGX agent

arXiv:2604.09037v1 Announce Type: cross Abstract: Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but over

Skip-Connected Policy Optimization for Implicit Advantage

Model ReleasesDGX agent

arXiv:2604.08690v1 Announce Type: cross Abstract: Group Relative Policy Optimization (GRPO) has proven effective in RLVR by using outcome-based rewards. While fine-grained dense rewards can theoretica

SPASM: Stable Persona-driven Agent Simulation for Multi-turn Dialogue Generation

Model ReleasesDGX agent

arXiv:2604.09212v1 Announce Type: new Abstract: Large language models are increasingly deployed in multi-turn settings such as tutoring, support, and counseling, where reliability depends on preservin

SSPO: Subsentence-level Policy Optimization

SafetyDGX agent

arXiv:2511.04256v2 Announce Type: replace Abstract: As a key component of large language model (LLM) post-training, Reinforcement Learning from Verifiable Rewards (RLVR) has substantially improved rea

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions

ResearchDGX agent

arXiv:2604.08477v1 Announce Type: cross Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has significantly improved large language model (LLM) reasoning in formal domains such as mathem

SynDocDis: A Metadata-Driven Framework for Generating Synthetic Physician Discussions Using Large Language Models

ApplicationsDGX agent

arXiv:2604.08555v1 Announce Type: new Abstract: Physician-physician discussions of patient cases represent a rich source of clinical knowledge and reasoning that could feed AI agents to enrich and eve

Task-Aware LLM Routing with Multi-Level Task-Profile-Guided Data Synthesis for Cold-Start Scenarios

ResearchDGX agent

arXiv:2604.09377v1 Announce Type: new Abstract: Large language models (LLMs) exhibit substantial variability in performance and computational cost across tasks and queries, motivating routing systems

Task Vectors, Learned Not Extracted: Performance Gains and Mechanistic Insight

ResearchDGX agent

arXiv:2509.24169v2 Announce Type: replace Abstract: Large Language Models (LLMs) can perform new tasks from in-context demonstrations, a phenomenon known as in-context learning (ICL). Recent work sugg

TaxPraBen: A Scalable Benchmark for Structured Evaluation of LLMs in Chinese Real-World Tax Practice

Model ReleasesDGX agent

arXiv:2604.08948v1 Announce Type: new Abstract: While Large Language Models (LLMs) excel in various general domains, they exhibit notable gaps in the highly specialized, knowledge-intensive, and legal

Testing the Assumptions of Active Learning for Translation Tasks with Few Samples

ResearchDGX agent

arXiv:2604.08977v1 Announce Type: new Abstract: Active learning (AL) is a training paradigm for selecting unlabeled samples for annotation to improve model performance on a test set, which is useful w

The Roots of Performance Disparity in Multilingual Language Models: Intrinsic Modeling Difficulty or Design Choices?

Model ReleasesDGX agent

arXiv:2601.07220v3 Announce Type: replace Abstract: Multilingual language models (LMs) promise broader NLP access, yet current systems deliver uneven performance across the world's languages. This sur

Think Less, Know More: State-Aware Reasoning Compression with Knowledge Guidance for Efficient Reasoning

SafetyDGX agent

arXiv:2604.09150v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) achieve strong performance on complex tasks by leveraging long Chain-of-Thought (CoT), but often suffer from overthinking,

UIPress: Bringing Optical Token Compression to UI-to-Code Generation

ResearchDGX agent

arXiv:2604.09442v1 Announce Type: new Abstract: UI-to-Code generation requires vision-language models (VLMs) to produce thousands of tokens of structured HTML/CSS from a single screenshot, making visu

Where Vision Becomes Text: Locating the OCR Routing Bottleneck in Vision-Language Models

Model ReleasesDGX agent

arXiv:2602.22918v2 Announce Type: replace Abstract: Vision-language models (VLMs) can read text from images, but where does this optical character recognition (OCR) information enter the language proc

Which Pieces Does Unigram Tokenization Really Need?

Model ReleasesDGX agent

arXiv:2512.12641v2 Announce Type: replace Abstract: The Unigram tokenization algorithm offers a probabilistic alternative to the greedy heuristics of Byte-Pair Encoding. Despite its theoretical elegan

You Can't Fight in Here! This is BBS!

TutorialsDGX agent

arXiv:2604.09501v1 Announce Type: new Abstract: Norm, the formal theoretical linguist, and Claudette, the computational language scientist, have a lovely time discussing whether modern language models

10 Apr 2026

A Decomposition Perspective to Long-context Reasoning for LLMs

ApplicationsDGX agent

arXiv:2604.07981v1 Announce Type: new Abstract: Long-context reasoning is essential for complex real-world applications, yet remains a significant challenge for Large Language Models (LLMs). Despite t

A GAN and LLM-Driven Data Augmentation Framework for Dynamic Linguistic Pattern Modeling in Chinese Sarcasm Detection

ResearchDGX agent

arXiv:2604.08381v1 Announce Type: new Abstract: Sarcasm is a rhetorical device that expresses criticism or emphasizes characteristics of certain individuals or situations through exaggeration, irony,

A systematic framework for generating novel experimental hypotheses from language models

SafetyDGX agent

arXiv:2408.05086v3 Announce Type: replace Abstract: Neural language models (LMs) have been shown to capture complex linguistic patterns, yet their utility in understanding human language and more broa

ACIArena: Toward Unified Evaluation for Agent Cascading Injection

Model ReleasesDGX agent

arXiv:2604.07775v1 Announce Type: cross Abstract: Collaboration and information sharing empower Multi-Agent Systems (MAS) but also introduce a critical security risk known as Agent Cascading Injection

ADAG: Automatically Describing Attribution Graphs

Model ReleasesDGX agent

arXiv:2604.07615v1 Announce Type: new Abstract: In language model interpretability research, extbf{circuit tracing} aims to identify which internal features causally contributed to a particular outp

Ads in AI Chatbots? An Analysis of How Large Language Models Navigate Conflicts of Interest

Model ReleasesDGX agent

arXiv:2604.08525v1 Announce Type: cross Abstract: Today's large language models (LLMs) are trained to align with user preferences through methods such as reinforcement learning. Yet models are beginni

AfriVoices-KE: A Multilingual Speech Dataset for Kenyan Languages

ResearchDGX agent

arXiv:2604.08448v1 Announce Type: new Abstract: AfriVoices-KE is a large-scale multilingual speech dataset comprising approximately 3,000 hours of audio across five Kenyan languages: Dholuo, Kikuyu, K

AI generates well-liked but templatic empathic responses

ApplicationsDGX agent

arXiv:2604.08479v1 Announce Type: new Abstract: Recent research shows that greater numbers of people are turning to Large Language Models (LLMs) for emotional support, and that people rate LLM respons

Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference

Model ReleasesDGX agent

arXiv:2604.08133v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) has become a dominant architecture for scaling large language models due to their sparse activation mechanism. However, the s

An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks

SafetyDGX agent

arXiv:2604.07883v1 Announce Type: cross Abstract: History textbooks often contain implicit biases, nationalist framing, and selective omissions that are difficult to audit at scale. We propose an agen

An Empirical Analysis of Static Analysis Methods for Detection and Mitigation of Code Library Hallucinations

ResearchDGX agent

arXiv:2604.07755v1 Announce Type: new Abstract: Despite extensive research, Large Language Models continue to hallucinate when generating code, particularly when using libraries. On NL-to-code benchma

Are GUI Agents Focused Enough? Automated Distraction via Semantic-level UI Element Injection

SafetyDGX agent

arXiv:2604.07831v1 Announce Type: cross Abstract: Existing red-teaming studies on GUI agents have important limitations. Adversarial perturbations typically require white-box access, which is unavaila

arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation

Model ReleasesDGX agent

arXiv:2504.10284v5 Announce Type: replace Abstract: Literature review tables are essential for summarizing and comparing collections of scientific papers. In this paper, we study the automatic generat

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention

ResearchDGX agent

arXiv:2604.07815v1 Announce Type: new Abstract: Long-context inference in LLMs faces the dual challenges of quadratic attention complexity and prohibitive KV cache memory. While token-level sparse att

AtomEval: Atomic Evaluation of Adversarial Claims in Fact Verification

ResearchDGX agent

arXiv:2604.07967v1 Announce Type: new Abstract: Adversarial claim rewriting is widely used to test fact-checking systems, but standard metrics fail to capture truth-conditional consistency and often l

Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test

Local AiDGX agent

arXiv:2506.06975v5 Announce Type: replace-cross Abstract: As API access becomes a primary interface to large language models (LLMs), users often interact with black-box systems that offer little trans

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

Model ReleasesDGX agent

arXiv:2604.08540v1 Announce Type: cross Abstract: Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchma

Behavior-Aware Item Modeling via Dynamic Procedural Solution Representations for Knowledge Tracing

ResearchDGX agent

arXiv:2604.08260v1 Announce Type: new Abstract: Knowledge Tracing (KT) aims to predict learners' future performance from past interactions. While recent KT approaches have improved via learning item r

BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity

Model ReleasesDGX agent

arXiv:2603.18019v2 Announce Type: replace Abstract: Do language model benchmarks actually measure what practitioners intend them to ? High-level metadata is too coarse to convey the granular reality o

Beyond Social Pressure: Benchmarking Epistemic Attack in Large Language Models

Model ReleasesDGX agent

arXiv:2604.07749v1 Announce Type: new Abstract: Large language models (LLMs) can shift their answers under pressure in ways that reflect accommodation rather than reasoning. Prior work on sycophancy h

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

Model ReleasesDGX agent

arXiv:2601.02670v2 Announce Type: replace Abstract: We introduce self-jailbreaking, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which oft

CAMO: A Class-Aware Minority-Optimized Ensemble for Robust Language Model Evaluation on Imbalanced Data

Model ReleasesDGX agent

arXiv:2604.07583v1 Announce Type: new Abstract: Real-world categorization is severely hampered by class imbalance because traditional ensembles favor majority classes, which lowers minority performanc

Can Vision Language Models Judge Action Quality? An Empirical Evaluation

Model ReleasesDGX agent

arXiv:2604.08294v1 Announce Type: cross Abstract: Action Quality Assessment (AQA) has broad applications in physical therapy, sports coaching, and competitive judging. Although Vision Language Models

Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization

Model ReleasesDGX agent

arXiv:2508.13993v2 Announce Type: replace Abstract: Long-context modeling is critical for a wide range of real-world tasks, including long-context question answering, summarization, and complex reason

ClawBench: Can AI Agents Complete Everyday Online Tasks?

Model ReleasesDGX agent

arXiv:2604.08523v1 Announce Type: new Abstract: AI agents may be able to automate your inbox, but can they automate other routine aspects of your life? Everyday online tasks offer a realistic yet unso

Clickbait detection: quick inference with maximum impact

ResearchDGX agent

arXiv:2604.08148v1 Announce Type: new Abstract: We propose a lightweight hybrid approach to clickbait detection that combines OpenAI semantic embeddings with six compact heuristic features capturing s

Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs

Local AiDGX agent

arXiv:2603.20698v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in medical image analysis. However, their application in gastr

Compact Example-Based Explanations for Language Models

ResearchDGX agent

arXiv:2601.03786v2 Announce Type: replace Abstract: Training data influence estimation methods quantify the contribution of training documents to a model's output, making them a promising source of in

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training

Model ReleasesDGX agent

arXiv:2604.07484v1 Announce Type: cross Abstract: Generative reward models (GRMs) have emerged as a promising approach for aligning Large Language Models (LLMs) with human preferences by offering grea

← Previous
1…123124125126127128
Next →