AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
13 May 2026

Unlocking LLM Creativity in Science through Analogical Reasoning

AgentsDGX agent

arXiv:2605.11258v1 Announce Type: cross Abstract: Autonomous science promises to augment scientific discovery, particularly in complex fields like biomedicine. However, this requires AI systems that c

VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference

Model ReleasesDGX agent

arXiv:2605.11334v1 Announce Type: cross Abstract: LLM-as-Judge systems are widely deployed for automated evaluation, yet practitioners lack reliable methods to know when a judge's verdict should be tr

Video-Language Understanding: A Survey from Model Architecture, Model Training, and Data Perspectives

ResearchDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2406.05615v4 Announce Type: replace Abstract: Humans use multiple senses to comprehend the environment. Vision and language are two of the most vital senses since they allow us to easily communi

What makes a word hard to learn? Modeling L1 influence on English vocabulary difficulty

TutorialsDGX agent

arXiv:2605.12281v1 Announce Type: new Abstract: What makes a word difficult to learn, and how does the difficulty depend on the learner's native language? We computationally model vocabulary difficult

When Emotion Becomes Trigger: Emotion-style dynamic Backdoor Attack Parasitising Large Language Models

ResearchDGX agent

arXiv:2605.11612v1 Announce Type: new Abstract: Backdoor vulnerabilities widely exist in the fine-tuning of large language models(LLMs). Most backdoor poisoning methods operate mainly at the token lev

When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated Content

ResearchDGX agent

arXiv:2512.17738v2 Announce Type: replace Abstract: User-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, ch

World Action Models: The Next Frontier in Embodied AI

SafetyDGX agent

arXiv:2605.12090v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have achieved strong semantic generalization for embodied policy learning, yet they learn reactive observation-to-

YFPO: A Preliminary Study of Yoked Feature Preference Optimization with Neuron-Guided Rewards for Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2605.11906v1 Announce Type: new Abstract: Preference optimization has become an important post-training paradigm for improving the reasoning abilities of large language models. Existing methods

12 May 2026

100,000+ Movie Reviews from Kazakhstan: Russian, Kazakh, and Code-Switched Texts

Model ReleasesDGX agent

arXiv:2605.08600v1 Announce Type: new Abstract: We present a new publicly available corpus of 100,502 movie reviews from Kazakhstan collected from kino.kz, spanning 2001-2025 and covering 4,943 unique

A Computational Operationalisation of Competing Maturational Theories of Syntactic Development via Statistical Grammar Induction

ResearchDGX agent

arXiv:2605.08476v1 Announce Type: new Abstract: This paper is concerned with what intermediate syntactic categories children acquire during first language development, and in what order. Maturational

A Single-Layer Model Can Do Language Modeling

ResearchDGX agent

arXiv:2605.10643v1 Announce Type: new Abstract: Modern language models scale depth by stacking layers, each holding its own state - a per-layer KV cache in transformers, a per-layer matrix in Mamba, G

A Single Layer to Explain Them All:Understanding Massive Activations in Large Language Models

ResearchDGX agent

arXiv:2605.08504v1 Announce Type: new Abstract: We investigate the origins of massive activations in large language models (LLMs) and identify a specific layer named the extbf{Massive Emergence Layer

AAAC: Activation-Aware Adaptive Codebooks for 4-bit LLM Weight Quantization

HardwareDGX agent

arXiv:2605.08692v1 Announce Type: cross Abstract: Post-training weight-only quantization to 4 bits is widely used to reduce the memory and compute costs of large language model inference. Existing PTQ

AdaSwitch: Adaptive Switching between Small and Large Agents for Effective Cloud-Local Collaborative Learning

Local AiDGX agent

arXiv:2410.13181v2 Announce Type: replace Abstract: Recent advancements in large language models (LLMs) have been remarkable. Users face a choice between using cloud-based LLMs for generation quality

AgentReview: Exploring Peer Review Dynamics with LLM Agents

SafetyDGX agent

arXiv:2406.12708v3 Announce Type: replace Abstract: Peer review is fundamental to the integrity and advancement of scientific publication. Traditional methods of peer review analyses often rely on exp

Aligning LLM Uncertainty with Human Disagreement in Subjectivity Analysis

SafetyDGX agent

arXiv:2605.10415v1 Announce Type: new Abstract: Large language models for subjectivity analysis are typically trained with aggregated labels, which compress variations in human judgment into a single

An Annotation Scheme and Classifier for Personal Facts in Dialogue

Model ReleasesDGX agent

arXiv:2605.10339v1 Announce Type: new Abstract: The advancement of Large Language Models (LLMs) has enabled their application in personalized dialogue systems. We present an extended annotation scheme

ANCHOR: Abductive Network Construction with Hierarchical Orchestration for Reliable Probability Inference in Large Language Models

ResearchDGX agent

arXiv:2605.10328v1 Announce Type: new Abstract: A central challenge in large-scale decision-making under incomplete information is estimating reliable probabilities. Recent approaches leverage Large L

Annotations Mitigate Post-Training Mode Collapse

TutorialsDGX agent

arXiv:2605.09995v1 Announce Type: new Abstract: Post-training (via supervised fine-tuning) improves instruction-following, but often induces semantic mode collapse by biasing models toward low-entropy

Architecture, Not Scale: Circuit Localization in Large Language Models

Model ReleasesDGX agent

arXiv:2605.08853v1 Announce Type: new Abstract: Mechanistic interpretability assumes that circuit analysis becomes harder as models scale. We challenge this assumption by showing that the attention ar

ASTRA-QA: A Benchmark for Abstract Question Answering over Documents

Model ReleasesDGX agent

arXiv:2605.10168v1 Announce Type: new Abstract: Document-based question answering (QA) increasingly includes abstract questions that require synthesizing scattered information from long documents or a

Attention Grounded Enhancement for Visual Document Retrieval

Model ReleasesDGX agent

arXiv:2511.13415v2 Announce Type: replace-cross Abstract: Visual document retrieval requires understanding heterogeneous and multi-modal content to satisfy implicit information needs. Recent advances

Autonomous Continual Learning for Environment Adaptation of Computer-Use Agents

AgentsDGX agent

arXiv:2602.10356v2 Announce Type: replace Abstract: Real-world digital environments are highly diverse and dynamic. These characteristics cause agents to frequently encounter unseen environments and d

BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation

Model ReleasesDGX agent

arXiv:2605.10845v1 Announce Type: cross Abstract: As global cross-lingual communication intensifies, language barriers in visually rich documents such as PDFs remain a practical bottleneck. Existing d

BaseCal: Unsupervised Confidence Calibration via Base Model Signals

ResearchDGX agent

arXiv:2601.03042v4 Announce Type: replace Abstract: Reliable confidence is essential for trusting the outputs of LLMs, yet widely deployed post-trained LLMs (PoLLMs) typically compromise this trust wi

BetaEdit: Null-Space Constrained Sequential Model Editing

ResearchDGX agent

arXiv:2605.09285v1 Announce Type: new Abstract: Null-space-based methods have garnered considerable attention in model editing by constraining updates to the null space of the pre-existing knowledge r

Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models

Model ReleasesDGX agent

arXiv:2605.09496v1 Announce Type: new Abstract: Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether

Beyond Local Edits: Embedding-Virtualized Knowledge for Broader Evaluation and Preservation of Model Editing

Model ReleasesDGX agent

arXiv:2602.01977v2 Announce Type: replace Abstract: Knowledge editing methods for large language models are commonly evaluated using predefined benchmarks that assess edited facts together with a limi

Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks

ResearchDGX agent

arXiv:2605.09955v1 Announce Type: new Abstract: Disagreement in annotation is a common phenomenon in the development of NLP datasets and serves as a valuable source of insight. While majority voting r

Beyond Multiple Choice: Evaluating Steering Vectors for Summarization

SafetyDGX agent

arXiv:2505.24859v3 Announce Type: replace-cross Abstract: Steering vectors are a lightweight method for controlling text properties by adding a learned bias to language model activations at inference

Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven

SafetyDGX agent

arXiv:2605.09463v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks. However, their deployment in long-context scenarios faces h

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence

Model ReleasesDGX agent

arXiv:2605.09041v1 Announce Type: new Abstract: Bias audits of large language models now operate within governance frameworks such as the EU AI Act, making benchmark reliability a security concern in

Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge

SafetyDGX agent

arXiv:2605.08123v1 Announce Type: cross Abstract: We study long-context balanced entropic optimal transport (OT) attention on TPU hardware through a stopped-base, fixed-depth tail-refinement surrogate

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

SafetyDGX agent

arXiv:2605.08721v1 Announce Type: new Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language game

Building Korean linguistic resource for NLU data generation of banking app CS dialog system

ResearchDGX agent

arXiv:2605.10241v1 Announce Type: new Abstract: Natural language understanding (NLU) is integral to task-oriented dialog systems, but demands a considerable amount of annotated training data to increa

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks

Model ReleasesDGX agent

arXiv:2605.09611v1 Announce Type: new Abstract: This preprint presents an empirical analysis of byte-exact chunk-level deduplication in Retrieval-Augmented Generation (RAG) pipelines. We measure conte

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges

ResearchDGX agent

arXiv:2605.09702v1 Announce Type: cross Abstract: Multi-judge evaluation is increasingly used to assess LLMs and reward models, and the prevailing heuristic is to curate: keep the most accurate judges

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

Model ReleasesDGX agent

arXiv:2601.12369v3 Announce Type: replace Abstract: Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing the

Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?

ResearchDGX agent

arXiv:2605.08439v1 Announce Type: new Abstract: Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where

Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness

Model ReleasesDGX agent

arXiv:2605.09634v1 Announce Type: new Abstract: LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability ac

cantnlp@DravidianLangTech 2026: organic domain adaptation improves multi-class hope speech detection in Tulu

ResearchDGX agent

arXiv:2605.09795v1 Announce Type: new Abstract: This paper presents our systems and results for the Hope Speech Detection in Code-Mixed Tulu Language shared task at the Sixth Workshop on Speech, Visio

Change My View? The Dynamics of Persuasion and Polarization in Online Discourse

SafetyDGX agent

arXiv:2605.08383v1 Announce Type: new Abstract: Philosophical accounts of persuasion often assume that shared evidence and rational argumentation should lead to a convergence of views between peers, y

Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus

Model ReleasesDGX agent

arXiv:2605.09092v1 Announce Type: new Abstract: This study addresses automatic transliteration from Tajik (Cyrillic script) to Persian (Perso-Arabic script). We present a curated, lexicographically ve

ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour

Model ReleasesDGX agent

arXiv:2506.12090v2 Announce Type: replace Abstract: This paper introduces ChatbotManip, a novel dataset for studying manipulation in Chatbots. It contains simulated generated conversations between a c

CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis

Model ReleasesDGX agent

arXiv:2510.11233v3 Announce Type: replace Abstract: Depression is a pressing global public health issue, yet publicly available Chinese-language resources for depression risk detection remain scarce a

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

SafetyDGX agent

arXiv:2602.11181v2 Announce Type: replace Abstract: Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling

Coherency through formalisations of Structured Natural Language, A case study on FRETish

TutorialsDGX agent

arXiv:2605.10462v1 Announce Type: new Abstract: Formalisation is the process of writing system requirements in a formal language. These requirements mostly originate in Natural Language. In the field

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

Model ReleasesDGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models

Model ReleasesDGX agent

arXiv:2601.16836v3 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to a

Complete Evidence Extraction with Model Ensembles: A Case Study on Medical Coding

ApplicationsDGX agent

arXiv:2511.07055v3 Announce Type: replace Abstract: High-stakes decisions informed by decision support systems require explicit evidence. While prior work focuses on short sufficient evidence, regulat

Composing Policy Gradients and Prompt Optimization for Language Model Programs

SafetyDGX agent

arXiv:2508.04660v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has proven to be an effective tool for post-training language models (LMs). However, AI systems are increa

Compute Where it Counts: Self Optimizing Language Models

SafetyDGX agent

arXiv:2605.10875v1 Announce Type: cross Abstract: Efficient LLM inference research has largely focused on reducing the cost of each decoding step (e.g., using quantization, pruning, or sparse attentio

ConFit v3: Improving Resume-Job Matching with LLM-based Re-Ranking

Model ReleasesDGX agent

arXiv:2605.09760v1 Announce Type: new Abstract: A reliable resume-job matching system helps a company find suitable candidates from a pool of resumes and helps a job seeker find relevant jobs from a l

Conformity Generates Collective Misalignment in AI Agents Societies

SafetyDGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation

Model ReleasesDGX agent

arXiv:2605.08522v1 Announce Type: new Abstract: The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequent

CREATE: Testing LLMs for Associative Creativity

Model ReleasesDGX agent

arXiv:2603.09970v2 Announce Type: replace Abstract: A key component of creativity is associative reasoning: the ability to draw novel yet meaningful connections between concepts. We introduce CREATE,

Cross-Cultural Transfer of Emoji Semantics and Sentiment in Financial Social Media

ResearchDGX agent

arXiv:2605.09414v1 Announce Type: new Abstract: Emojis are widely used in online financial communication, but it is unclear whether they provide transferable sentiment signals across languages, platfo

Crosslingual On-Policy Self-Distillation for Multilingual Reasoning

SafetyDGX agent

arXiv:2605.09548v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress in mathematical reasoning, but this ability is not equally accessible across languages. E

DECO-MWE: building a linguistic resource of Korean multiword expressions for feature-based sentiment analysis

ResearchDGX agent

arXiv:2605.10295v1 Announce Type: new Abstract: This paper aims to construct a linguistic resource of Korean Multiword Expressions for Feature-Based Sentiment Analysis (FBSA): DECO-MWE. Dealing with m

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

Model ReleasesDGX agent

arXiv:2605.10933v1 Announce Type: cross Abstract: While Mixture-of-Experts (MoE) scales model capacity without proportionally increasing computation, its massive total parameter footprint creates sign

← Previous
1…7778798081…129
Next →