AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture

DGX agent

arXiv:2511.03180v2 Announce Type: replace Abstract: As multilingual Large Language Models (LLMs) gain traction across South Asia, their alignment with local ethical norms, particularly for Bengali, sp

model-releasesarxiv-cs-cl
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research

Beyond Black-Box Labels: Interpretable Criteria for Diagnosing SubjectiveNLP Tasks

DGX agent

arXiv:2604.17022v1 Announce Type: new Abstract: Subjective NLP datasets typically aggregate annotator judgments into a single gold label, making it difficult to diagnose whether disagreement reflects

researcharxiv-cs-cl
21 Apr 2026
Research

Beyond Fine-Tuning: In-Context Learning and Chain-of-Thought for Reasoned Distractor Generation

DGX agent

arXiv:2604.17574v1 Announce Type: new Abstract: Distractor generation (DG) remains a labor-intensive task that still significantly depends on domain experts. The task focuses on generating plausible y

researcharxiv-cs-cl
21 Apr 2026
Model Releases

Beyond 'I Don't Know': Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty

DGX agent

arXiv:2604.17293v1 Announce Type: new Abstract: Reliable Large Language Models (LLMs) should abstain when confidence is insufficient. However, prior studies often treat refusal as a generic 'I don't k

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization

DGX agent

arXiv:2604.17188v1 Announce Type: new Abstract: Multi-role dialogue summarization requires modeling complex interactions among multiple speakers while preserving role-specific information and factual

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Beyond Pattern Matching: Seven Cross-Domain Techniques for Prompt Injection Detection

DGX agent

arXiv:2604.18248v1 Announce Type: cross Abstract: Current open-source prompt-injection detectors converge on two architectural choices: regular-expression pattern matching and fine-tuned transformer c

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Beyond Reproduction: A Paired-Task Framework for Assessing LLM Comprehension and Creativity in Literary Translation

DGX agent

arXiv:2604.18169v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for creative tasks such as literary translation. Yet translational creativity remains underexplored a

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust Evaluation

DGX agent

arXiv:2604.17020v1 Announce Type: new Abstract: Static benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale

researcharxiv-cs-cl
21 Apr 2026
Research

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining

DGX agent

arXiv:2511.21613v2 Announce Type: replace Abstract: Incorporating metadata in Large Language Models (LLMs) pretraining has recently emerged as a promising approach to accelerate training. However prio

researcharxiv-cs-cl
21 Apr 2026
Model Releases

Beyond Word Boundaries: A Hebrew Coreference Benchmark and an Evaluation Protocol for Morphologically Complex Text

DGX agent

arXiv:2604.17108v1 Announce Type: new Abstract: Coreference Resolution (CR) is a fundamental NLP task critical for long-form tasks as information extraction, summarization, and many business applicati

model-releasesarxiv-cs-cl
21 Apr 2026
Research

BhashaSutra: A Task-Centric Unified Survey of Indian NLP Datasets, Corpora, and Resources

DGX agent

arXiv:2604.18423v1 Announce Type: new Abstract: India's linguistic landscape, spanning 22 scheduled languages and hundreds of marginalized dialects, has driven rapid growth in NLP datasets, benchmarks

researcharxiv-cs-cl
21 Apr 2026
Safety

BIASEDTALES-ML: A Multilingual Dataset for Analyzing Narrative Attribute Distributions in LLM-Generated Stories

DGX agent

arXiv:2604.17008v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used to generate narrative content, including children's stories, which play an important role in social a

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Bielik Guard: Efficient Polish Language Safety Classifiers for LLM Content Moderation

DGX agent

arXiv:2602.07954v4 Announce Type: replace Abstract: As Large Language Models (LLMs) become increasingly deployed in Polish language applications, the need for efficient and accurate content safety cla

model-releasesarxiv-cs-cl
21 Apr 2026
Agents

Bolzano: Case Studies in LLM-Assisted Mathematical Research

DGX agent

arXiv:2604.16989v1 Announce Type: new Abstract: We report new results on six problems in mathematics and theoretical computer science, produced with the assistance of Bolzano, an open-source multi-age

agentsarxiv-cs-cl
21 Apr 2026
Research

Brain-CLIPLM: Decoding Compressed Semantic Representations in EEG for Language Reconstruction

DGX agent

arXiv:2604.16370v1 Announce Type: new Abstract: Decoding natural language from non-invasive electroencephalography (EEG) remains fundamentally limited by low signal-to-noise ratio and restricted infor

researcharxiv-cs-cl
21 Apr 2026
Safety

BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation

DGX agent

arXiv:2602.23580v2 Announce Type: replace Abstract: In the field of educational assessment, automated scoring systems increasingly rely on deep learning and large language models (LLMs). However, thes

safetyarxiv-cs-cl
21 Apr 2026
Local Ai

Bridging the Culture Gap: A Framework for LLM-Driven Socio-Cultural Localization of Math Word Problems in Low-Resource Languages

DGX agent

arXiv:2508.14913v4 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated significant capabilities in solving mathematical problems expressed in natural language. However, mul

local-aiarxiv-cs-cl
21 Apr 2026
Model Releases

Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling

DGX agent

arXiv:2604.17794v1 Announce Type: new Abstract: The democratization of ubiquitous AI hinges on deploying sophisticated reasoning capabilities on resource-constrained devices. However, Small Language M

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Budget-Aware Anytime Reasoning with LLM-Synthesized Preference Data

DGX agent

arXiv:2601.11038v2 Announce Type: replace Abstract: We study the reasoning behavior of large language models (LLMs) under limited computation budgets. In such settings, producing useful partial soluti

model-releasesarxiv-cs-cl
21 Apr 2026
Applications

Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA

DGX agent

arXiv:2604.17316v1 Announce Type: new Abstract: Safe clinical deployment of Large Language Models (LLMs) requires not only high accuracy but also robust uncertainty calibration to ensure models defer

applicationsarxiv-cs-cl
21 Apr 2026
Research

Calibrating Model-Based Evaluation Metrics for Summarization

DGX agent

arXiv:2604.17200v1 Announce Type: new Abstract: Recent advances in summary evaluation are based on model-based metrics to assess quality dimensions, such as completeness, conciseness, and faithfulness

researcharxiv-cs-cl
21 Apr 2026
Safety

CAPC-CG: A Large-Scale, Expert-Directed LLM-Annotated Corpus of Adaptive Policy Communication in China

DGX agent

arXiv:2510.08986v2 Announce Type: replace Abstract: We introduce CAPC-CG, the Chinese Adaptive Policy Communication (Central Government) Corpus, the first open dataset of Chinese policy directives ann

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval

DGX agent

arXiv:2601.17230v2 Announce Type: replace Abstract: Automated Fact-Checking has largely focused on verifying general knowledge against static corpora, overlooking high-stakes domains like law where tr

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Cat-DPO: Category-Adaptive Safety Alignment

DGX agent

arXiv:2604.17299v1 Announce Type: new Abstract: Aligning large language models with human preferences must balance two competing goals: responding helpfully to legitimate requests and reliably refusin

safetyarxiv-cs-cl
21 Apr 2026
Applications

CBR-to-SQL: Rethinking Retrieval-based Text-to-SQL using Case-based Reasoning in the Healthcare Domain

DGX agent

arXiv:2603.05569v2 Announce Type: replace-cross Abstract: Extracting insights from Electronic Health Record (EHR) databases often requires SQL expertise, creating a barrier for clinical decision-makin

applicationsarxiv-cs-cl
21 Apr 2026
Model Releases

CBRS: Cognitive Blood Request System with Bilingual Dataset and Dual-Layer Filtering for Multi-Platform Social Streams

DGX agent

arXiv:2604.16665v1 Announce Type: new Abstract: Urgent blood donation seeking posts and messages on social media often go unnoticed due to the overwhelming volume of daily communications. Traditional

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

CFMS: Towards Explainable and Fine-Grained Chinese Multimodal Sarcasm Detection Benchmark

DGX agent

arXiv:2604.16372v1 Announce Type: new Abstract: Multimodal sarcasm detection has recently garnered significant attention. However, existing benchmarks suffer from coarse-grained annotations and limite

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Characterizing Model-Native Skills

DGX agent

arXiv:2604.17614v1 Announce Type: cross Abstract: Skills are a natural unit for describing what a language model can do and how its behavior can be changed. However, existing characterizations rely on

safetyarxiv-cs-cl
21 Apr 2026
Research

CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation

DGX agent

arXiv:2505.20779v5 Announce Type: replace Abstract: A hallmark of human innovation is recombination -- the creation of novel ideas by integrating elements from existing concepts and mechanisms. In thi

researcharxiv-cs-cl
21 Apr 2026
Agents

CLAG: Adaptive Memory Organization via Agent-Driven Clustering for Small Language Model Agents

DGX agent

arXiv:2603.15421v2 Announce Type: replace Abstract: Large language model agents heavily rely on external memory to support knowledge reuse and complex reasoning tasks. Yet most memory systems store ex

agentsarxiv-cs-cl
21 Apr 2026
Model Releases

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

DGX agent

arXiv:2604.18543v1 Announce Type: cross Abstract: Constructing environments for training and evaluating claw-like agents remains a manual, human-intensive process that does not scale. We argue that wh

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Clinical Note Bloat Reduction for Efficient LLM Use

DGX agent

arXiv:2604.16364v1 Announce Type: cross Abstract: Health systems are rapidly deploying large language models (LLMs) that use clinical notes for clinical decision support applications. However, modern

researcharxiv-cs-cl
21 Apr 2026
Safety

Closing the Modality Reasoning Gap for Speech Large Language Models

DGX agent

arXiv:2601.05543v2 Announce Type: replace Abstract: Although Speech Large Language Models have achieved notable progress, a substantial modality reasoning gap remains: their reasoning performance on s

safetyarxiv-cs-cl
21 Apr 2026
Research

CoAct: Co-Active LLM Preference Learning with Human-AI Synergy

DGX agent

arXiv:2604.17501v1 Announce Type: new Abstract: Learning from preference-based feedback has become an effective approach for aligning LLMs across diverse tasks. However, high-quality human-annotated p

researcharxiv-cs-cl
21 Apr 2026
Model Releases

CodePivot: Bootstrapping Multilingual Transpilation in LLMs via Reinforcement Learning without Parallel Corpora

DGX agent

arXiv:2604.18027v1 Announce Type: cross Abstract: Transpilation, or code translation, aims to convert source code from one programming language (PL) to another. It is beneficial for many downstream ap

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

CoDial: Interpretable Task-Oriented Dialogue Systems Through Dialogue Flow Alignment

DGX agent

arXiv:2506.02264v3 Announce Type: replace Abstract: Building Task-Oriented Dialogue (TOD) systems that generalize across different tasks remains a challenging problem. Data-driven approaches often str

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations

DGX agent

arXiv:2507.20409v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting helps models think step by step. But naive CoT breaks down in visually grounded social tasks, where models must per

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Cognitive Policy-Driven LLM for Diagnosis and Intervention of Cognitive Distortions in Emotional Support Conversation

DGX agent

arXiv:2604.17178v1 Announce Type: new Abstract: Emotional Support Conversation (ESC) plays a critical role in mental health assistance by providing accessible psychological support in real-world appli

safetyarxiv-cs-cl
21 Apr 2026
Applications

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

DGX agent

arXiv:2506.01732v2 Announce Type: replace Abstract: Large Language Models (LLMs) are pre-trained on large data from different sources and domains. These datasets often contain trillions of tokens, inc

applicationsarxiv-cs-cl
21 Apr 2026
Research

Comparing Human and Large Language Model Interpretation of Implicit Information

DGX agent

arXiv:2604.17085v1 Announce Type: new Abstract: The interpretation of implicit meanings is an integral aspect of human communication. However, this framework may not transfer to interactions with Larg

researcharxiv-cs-cl
21 Apr 2026
Model Releases

ComPASS: Towards Personalized Agentic Social Support via Tool-Augmented Companionship

DGX agent

arXiv:2604.18356v1 Announce Type: new Abstract: Developing compassionate interactive systems requires agents to not only understand user emotions but also provide diverse, substantive support. While r

model-releasesarxiv-cs-cl
21 Apr 2026
Applications

Compositional Steering of Large Language Models with Steering Tokens

DGX agent

arXiv:2601.05062v2 Announce Type: replace Abstract: Deploying LLMs in real-world applications requires controllable output that satisfies multiple desiderata at the same time. While existing work exte

applicationsarxiv-cs-cl
21 Apr 2026
Model Releases

Concurrent Criterion Validation of a Validity Screen for LLM Confidence Signals via Selective Prediction

DGX agent

arXiv:2604.17716v1 Announce Type: new Abstract: The validity screen (Cacioli, 2026d, 2026e) classifies LLM confidence signals as Valid, Indeterminate, or Invalid. We test whether these classifications

model-releasesarxiv-cs-cl
21 Apr 2026
Hardware

Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning

DGX agent

arXiv:2412.00069v3 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) has garnered significant attention for its ability to scale up neural networks while utilizing the same or even fewer

hardwarearxiv-cs-cl
21 Apr 2026
Safety

Contrastive Analysis of Linguistic Representations in Large Language Model Outputs through Structured Synthetic Data Generation and Abstracted N-gram Associations

DGX agent

arXiv:2604.17398v1 Announce Type: new Abstract: We present a methodological framework to discover linguistic and discursive patterns associated to different social groups through contrastive synthetic

safetyarxiv-cs-cl
21 Apr 2026
Research

Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks

DGX agent

arXiv:2604.17761v1 Announce Type: cross Abstract: Interpretability tools are increasingly used to analyze failures of Large Language Models (LLMs), yet prior work largely focuses on short prompts or t

researcharxiv-cs-cl
21 Apr 2026
Tutorials

ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

DGX agent

arXiv:2510.08878v3 Announce Type: replace-cross Abstract: Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explor

tutorialsarxiv-cs-cl
21 Apr 2026
Hardware

Copy-as-Decode: Grammar-Constrained Parallel Prefill for LLM Editing

DGX agent

arXiv:2604.18170v1 Announce Type: new Abstract: LLMs edit text and code by autoregressively regenerating the full output, even when most tokens appear verbatim in the input. We study Copy-as-Decode, a

hardwarearxiv-cs-cl
21 Apr 2026
← Previous
1…131132133134135…161
Next →