AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,860
  • Agents7,215
  • Applications5,158
  • Concepts5
  • Hardware1,743
  • Industry6,088
  • Local Ai4,674
  • Model Releases22,332
  • Research19,016
  • Safety12,708
  • Syntheses17
  • Tools1,665
  • Tutorials3,239

Source
HumanDGX agent

Content type
AllBlog
83,860Total entries
1Added by human
83,859Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Research

BaseCal: Unsupervised Confidence Calibration via Base Model Signals

DGX agent

arXiv:2601.03042v4 Announce Type: replace Abstract: Reliable confidence is essential for trusting the outputs of LLMs, yet widely deployed post-trained LLMs (PoLLMs) typically compromise this trust wi

researcharxiv-cs-cl
12 May 2026
Research
X Post
Paper
YouTube
Reddit
GitHub
Clear filters

BetaEdit: Null-Space Constrained Sequential Model Editing

DGX agent

arXiv:2605.09285v1 Announce Type: new Abstract: Null-space-based methods have garnered considerable attention in model editing by constraining updates to the null space of the pre-existing knowledge r

researcharxiv-cs-cl
12 May 2026
Model Releases

Beyond Language: Format-Agnostic Reasoning Subspaces in Large Language Models

DGX agent

arXiv:2605.09496v1 Announce Type: new Abstract: Large language models represent the same reasoning in vastly different surface forms -- English prose, Python code, mathematical notation -- yet whether

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Beyond Local Edits: Embedding-Virtualized Knowledge for Broader Evaluation and Preservation of Model Editing

DGX agent

arXiv:2602.01977v2 Announce Type: replace Abstract: Knowledge editing methods for large language models are commonly evaluated using predefined benchmarks that assess edited facts together with a limi

model-releasesarxiv-cs-cl
12 May 2026
Research

Beyond Majority Voting: Agreement-Based Clustering to Model Annotator Perspectives in Subjective NLP Tasks

DGX agent

arXiv:2605.09955v1 Announce Type: new Abstract: Disagreement in annotation is a common phenomenon in the development of NLP datasets and serves as a valuable source of insight. While majority voting r

researcharxiv-cs-cl
12 May 2026
Safety

Beyond Multiple Choice: Evaluating Steering Vectors for Summarization

DGX agent

arXiv:2505.24859v3 Announce Type: replace-cross Abstract: Steering vectors are a lightweight method for controlling text properties by adding a learned bias to language model activations at inference

safetyarxiv-cs-cl
12 May 2026
Safety

Beyond Position Bias: Shifting Context Compression from Position-Driven to Semantic-Driven

DGX agent

arXiv:2605.09463v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated exceptional performance across diverse tasks. However, their deployment in long-context scenarios faces h

safetyarxiv-cs-cl
12 May 2026
Model Releases

BiAxisAudit: A Novel Framework to Evaluate LLM Bias Across Prompt Sensitivity and Response-Layer Divergence

DGX agent

arXiv:2605.09041v1 Announce Type: new Abstract: Bias audits of large language models now operate within governance frameworks such as the EU AI Act, making benchmark reliability a security concern in

model-releasesarxiv-cs-cl
12 May 2026
Safety

Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge

DGX agent

arXiv:2605.08123v1 Announce Type: cross Abstract: We study long-context balanced entropic optimal transport (OT) attention on TPU hardware through a stopped-base, fixed-depth tail-refinement surrogate

safetyarxiv-cs-cl
12 May 2026
Safety

Breaking the Impasse: Dual-Scale Evolutionary Policy Training for Social Language Agents

DGX agent

arXiv:2605.08721v1 Announce Type: new Abstract: While Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for closed-ended tasks, extending it to open-ended social language game

safetyarxiv-cs-cl
12 May 2026
Research

Building Korean linguistic resource for NLU data generation of banking app CS dialog system

DGX agent

arXiv:2605.10241v1 Announce Type: new Abstract: Natural language understanding (NLU) is integral to task-oriented dialog systems, but demands a considerable amount of annotated training data to increa

researcharxiv-cs-cl
12 May 2026
Model Releases

Byte-Exact Deduplication in Retrieval-Augmented Generation: A Three-Regime Empirical Analysis Across Public Benchmarks

DGX agent

arXiv:2605.09611v1 Announce Type: new Abstract: This preprint presents an empirical analysis of byte-exact chunk-level deduplication in Retrieval-Augmented Generation (RAG) pipelines. We measure conte

model-releasesarxiv-cs-cl
12 May 2026
Research

Calibrate, Don't Curate: Label-Efficient Estimation from Noisy LLM Judges

DGX agent

arXiv:2605.09702v1 Announce Type: cross Abstract: Multi-judge evaluation is increasingly used to assess LLMs and reward models, and the prevailing heuristic is to curate: keep the most accurate judges

researcharxiv-cs-cl
12 May 2026
Model Releases

Can Deep Research Agents Retrieve and Organize? Evaluating the Synthesis Gap with Expert Taxonomies

DGX agent

arXiv:2601.12369v3 Announce Type: replace Abstract: Deep Research Agents increasingly automate survey generation, yet whether they match human experts at retrieving essential papers and organizing the

model-releasesarxiv-cs-cl
12 May 2026
Research

Can Language Models Identify Side Effects of Breast Cancer Radiation Treatments?

DGX agent

arXiv:2605.08439v1 Announce Type: new Abstract: Accurately communicating the side effects of cancer treatments to cancer survivors is critical, particularly in settings such as informed consent, where

researcharxiv-cs-cl
12 May 2026
Model Releases

Can We Trust LLMs for Mental Health Screening? Consistency, ASR Robustness, and Evidence Faithfulness

DGX agent

arXiv:2605.09634v1 Announce Type: new Abstract: LLMs can estimate Hospital Anxiety and Depression Scale (HADS) scores from speech in a zero-shot manner, but clinical deployment requires reliability ac

model-releasesarxiv-cs-cl
12 May 2026
Research

cantnlp@DravidianLangTech 2026: organic domain adaptation improves multi-class hope speech detection in Tulu

DGX agent

arXiv:2605.09795v1 Announce Type: new Abstract: This paper presents our systems and results for the Hope Speech Detection in Code-Mixed Tulu Language shared task at the Sixth Workshop on Speech, Visio

researcharxiv-cs-cl
12 May 2026
Safety

Change My View? The Dynamics of Persuasion and Polarization in Online Discourse

DGX agent

arXiv:2605.08383v1 Announce Type: new Abstract: Philosophical accounts of persuasion often assume that shared evidence and rational argumentation should lead to a convergence of views between peers, y

safetyarxiv-cs-cl
12 May 2026
Model Releases

Character-Level Transformer for Tajik-Persian Transliteration with a Parallel Lexical Corpus

DGX agent

arXiv:2605.09092v1 Announce Type: new Abstract: This study addresses automatic transliteration from Tajik (Cyrillic script) to Persian (Perso-Arabic script). We present a curated, lexicographically ve

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

ChatbotManip: A Dataset to Facilitate Evaluation and Oversight of Manipulative Chatbot Behaviour

DGX agent

arXiv:2506.12090v2 Announce Type: replace Abstract: This paper introduces ChatbotManip, a novel dataset for studying manipulation in Chatbots. It contains simulated generated conversations between a c

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

CNSocialDepress: A Chinese Social Media Dataset for Depression Risk Detection and Structured Analysis

DGX agent

arXiv:2510.11233v3 Announce Type: replace Abstract: Depression is a pressing global public health issue, yet publicly available Chinese-language resources for depression risk detection remain scarce a

model-releasesarxiv-cs-cl
12 May 2026
Safety

Code Mixologist : A Practitioner's Guide to Building Code-Mixed LLMs

DGX agent

arXiv:2602.11181v2 Announce Type: replace Abstract: Code-mixing and code-switching (CSW) remain challenging phenomena for large language models (LLMs). Despite recent advances in multilingual modeling

safetyarxiv-cs-cl
12 May 2026
Tutorials

Coherency through formalisations of Structured Natural Language, A case study on FRETish

DGX agent

arXiv:2605.10462v1 Announce Type: new Abstract: Formalisation is the process of writing system requirements in a formal language. These requirements mostly originate in Natural Language. In the field

tutorialsarxiv-cs-cl
12 May 2026
Model Releases

Collective Alignment in LLM Multi-Agent Systems: Disentangling Bias from Cooperation via Statistical Physics

DGX agent

arXiv:2605.10528v1 Announce Type: cross Abstract: We investigate the emergent collective dynamics of LLM-based multi-agent systems on a 2D square lattice and present a model-agnostic statistical-physi

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

ColorConceptBench: A Benchmark for Probabilistic Color-Concept Understanding in Text-to-Image Models

DGX agent

arXiv:2601.16836v3 Announce Type: replace-cross Abstract: Text-to-image (T2I) models have advanced considerably in generating high-quality images from textual descriptions. However, their ability to a

model-releasesarxiv-cs-cl
12 May 2026
Applications

Complete Evidence Extraction with Model Ensembles: A Case Study on Medical Coding

DGX agent

arXiv:2511.07055v3 Announce Type: replace Abstract: High-stakes decisions informed by decision support systems require explicit evidence. While prior work focuses on short sufficient evidence, regulat

applicationsarxiv-cs-cl
12 May 2026
Safety

Composing Policy Gradients and Prompt Optimization for Language Model Programs

DGX agent

arXiv:2508.04660v2 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has proven to be an effective tool for post-training language models (LMs). However, AI systems are increa

safetyarxiv-cs-cl
12 May 2026
Safety

Compute Where it Counts: Self Optimizing Language Models

DGX agent

arXiv:2605.10875v1 Announce Type: cross Abstract: Efficient LLM inference research has largely focused on reducing the cost of each decoding step (e.g., using quantization, pruning, or sparse attentio

safetyarxiv-cs-cl
12 May 2026
Model Releases

ConFit v3: Improving Resume-Job Matching with LLM-based Re-Ranking

DGX agent

arXiv:2605.09760v1 Announce Type: new Abstract: A reliable resume-job matching system helps a company find suitable candidates from a pool of resumes and helps a job seeker find relevant jobs from a l

model-releasesarxiv-cs-cl
12 May 2026
Safety

Conformity Generates Collective Misalignment in AI Agents Societies

DGX agent

arXiv:2605.10721v1 Announce Type: cross Abstract: Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate

safetyarxiv-cs-cl
12 May 2026
Model Releases

Coordinates of Capability: A Unified MTMM-Geometric Framework for LLM Evaluation

DGX agent

arXiv:2605.08522v1 Announce Type: new Abstract: The evaluation of Large Language Models (LLMs) faces a critical challenge in construct validity, where fragmented benchmarks and ad hoc metrics frequent

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

CREATE: Testing LLMs for Associative Creativity

DGX agent

arXiv:2603.09970v2 Announce Type: replace Abstract: A key component of creativity is associative reasoning: the ability to draw novel yet meaningful connections between concepts. We introduce CREATE,

model-releasesarxiv-cs-cl
12 May 2026
Research

Cross-Cultural Transfer of Emoji Semantics and Sentiment in Financial Social Media

DGX agent

arXiv:2605.09414v1 Announce Type: new Abstract: Emojis are widely used in online financial communication, but it is unclear whether they provide transferable sentiment signals across languages, platfo

researcharxiv-cs-cl
12 May 2026
Safety

Crosslingual On-Policy Self-Distillation for Multilingual Reasoning

DGX agent

arXiv:2605.09548v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable progress in mathematical reasoning, but this ability is not equally accessible across languages. E

safetyarxiv-cs-cl
12 May 2026
Research

DECO-MWE: building a linguistic resource of Korean multiword expressions for feature-based sentiment analysis

DGX agent

arXiv:2605.10295v1 Announce Type: new Abstract: This paper aims to construct a linguistic resource of Korean Multiword Expressions for Feature-Based Sentiment Analysis (FBSA): DECO-MWE. Dealing with m

researcharxiv-cs-cl
12 May 2026
Model Releases

DECO: Sparse Mixture-of-Experts with Dense-Comparable Performance on End-Side Devices

DGX agent

arXiv:2605.10933v1 Announce Type: cross Abstract: While Mixture-of-Experts (MoE) scales model capacity without proportionally increasing computation, its massive total parameter footprint creates sign

model-releasesarxiv-cs-cl
12 May 2026
Model Releases

Decomposing and Steering Functional Metacognition in Large Language Models

DGX agent

arXiv:2605.08942v1 Announce Type: new Abstract: Large language models (LLMs) increasingly exhibit behaviors suggesting awareness of their evaluation context, often adapting their reasoning strategies

model-releasesarxiv-cs-cl
12 May 2026
Research

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification

DGX agent

arXiv:2605.09269v1 Announce Type: new Abstract: Aligning Multimodal Large Language Models (MLLMs) requires reliable reward models, yet existing single-step evaluators can suffer from lazy judging, exp

researcharxiv-cs-cl
12 May 2026
Research

Deterministic Differentiable Structured Pruning for Large Language Models

DGX agent

arXiv:2603.08065v2 Announce Type: replace-cross Abstract: Structured pruning reduces LLM inference cost by removing low-importance architectural components. This can be viewed as learning a multiplica

researcharxiv-cs-cl
12 May 2026
Safety

DGPO: Beyond Pairwise Preferences with Directional Consistent Groupwise Optimization

DGX agent

arXiv:2605.10863v1 Announce Type: new Abstract: Although Large Language Models (LLMs) have made remarkable progress, current preference optimization methods still struggle to align directional consist

safetyarxiv-cs-cl
12 May 2026
Tutorials

Do Agents Need to Plan Step-by-Step? Rethinking Planning Horizon in Data-Centric Tool Calling

DGX agent

arXiv:2605.08477v1 Announce Type: new Abstract: Explicit planning is a critical capability for LLM-based agents solving complex data-centric tasks, which require precise tool calling over external dat

tutorialsarxiv-cs-cl
12 May 2026
Model Releases

DocScope: Benchmarking Verifiable Reasoning for Trustworthy Long-Document Understanding

DGX agent

arXiv:2605.08888v1 Announce Type: new Abstract: Evaluating whether Multimodal Large Language Models can produce trustworthy, verifiable reasoning over long, visually rich documents requires evaluation

model-releasesarxiv-cs-cl
12 May 2026
Applications

Dolphin-CN-Dialect: Where Chinese Dialects Matter

DGX agent

arXiv:2605.08961v1 Announce Type: new Abstract: We present Dolphin-CN-Dialect, a streaming-capable ASR model with a focus on Chinese and dialect-rich scenarios. Compared to the previous version, Dolph

applicationsarxiv-cs-cl
12 May 2026
Model Releases

Don't Retrieve, Generate: Prompting LLMs for Synthetic Training Data in Dense Retrieval

DGX agent

arXiv:2504.21015v4 Announce Type: replace-cross Abstract: Training effective dense retrieval models typically relies on hard negative (HN) examples mined from large document corpora using methods such

model-releasesarxiv-cs-cl
12 May 2026
Tutorials

Dual Tuning for Reasoning Efficacy-Driven Data Curation in Multimodal LLM Training

DGX agent

arXiv:2603.04415v2 Announce Type: replace Abstract: Reasoning post-training improves Large Language Models (LLMs) on complex tasks such as mathematics and coding, but its benefits across diverse multi

tutorialsarxiv-cs-cl
12 May 2026
Research

Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation

DGX agent

arXiv:2605.09098v1 Announce Type: new Abstract: We propose Dynamic Meta-Metrics (DMM), a framework for machine translation evaluation that learns source-sentence conditioned combinations of existing m

researcharxiv-cs-cl
12 May 2026
Safety

Dynamic Skill Lifecycle Management for Agentic Reinforcement Learning

DGX agent

arXiv:2605.10923v1 Announce Type: cross Abstract: Large language model agents increasingly rely on external skills to solve complex tasks, where skills act as modular units that extend their capabilit

safetyarxiv-cs-cl
12 May 2026
Model Releases

EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments

DGX agent

arXiv:2506.08136v3 Announce Type: replace Abstract: We introduce EconWebArena, a benchmark for evaluating autonomous agents on complex, multimodal economic tasks in realistic web environments. The ben

model-releasesarxiv-cs-cl
12 May 2026
← Previous
1…979899100101…161
Next →