AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Hardware

PiERN: Token-Level Routing for Integrating High-Precision Computation and Reasoning

DGX agent

arXiv:2509.18169v3 Announce Type: replace-cross Abstract: Tasks on complex systems require high-precision numerical computation to support decisions, but current large language models (LLMs) cannot in

hardwarearxiv-cs-cl
21 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tutorials

Plausibility as Commonsense Reasoning: Humans Succeed, Large Language Models Do not

DGX agent

arXiv:2604.04825v2 Announce Type: replace Abstract: Large language models achieve strong performance on many language tasks, yet it remains unclear whether they integrate world knowledge with syntacti

tutorialsarxiv-cs-cl
21 Apr 2026
Model Releases

Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding

DGX agent

arXiv:2604.17132v1 Announce Type: new Abstract: Safety-aligned large language models (LLMs) often generate refusal responses to harmless queries due to the over-refusal problem. However, existing meth

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

PoliLegalLM: A Technical Report on a Large Language Model for Political and Legal Affairs

DGX agent

arXiv:2604.17543v1 Announce Type: new Abstract: Large language models (LLMs) have achieved remarkable success in general-domain tasks, yet their direct application to the legal domain remains challeng

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Polysemantic Experts, Monosemantic Paths: Routing as Control in MoEs

DGX agent

arXiv:2604.17837v1 Announce Type: cross Abstract: An LLM's residual stream is both state and instruction: it encodes the current context and determines the next transformation. We introduce a paramete

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning

DGX agent

arXiv:2502.02871v2 Announce Type: replace Abstract: Scientific reasoning, the process through which humans apply logic, evidence, and critical thinking to explore and interpret scientific phenomena, i

researcharxiv-cs-cl
21 Apr 2026
Model Releases

Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?

DGX agent

arXiv:2604.17338v1 Announce Type: cross Abstract: Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often regenerate correct but o

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

PrefixMemory-Tuning: Modernizing Prefix-Tuning by Decoupling the Prefix from Attention

DGX agent

arXiv:2506.13674v3 Announce Type: replace Abstract: Parameter-Efficient Fine-Tuning (PEFT) methods have become crucial for rapidly adapting large language models (LLMs) to downstream tasks. Prefix-Tun

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

PrinciplismQA: A Philosophy-Grounded Approach to Assessing LLM-Human Clinical Medical Ethics Alignment

DGX agent

arXiv:2508.05132v2 Announce Type: replace Abstract: As medical LLMs transition to clinical deployment, assessing their ethical reasoning capability becomes critical. While achieving high accuracy on k

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

PRISM: Probing Reasoning, Instruction, and Source Memory in LLM Hallucinations

DGX agent

arXiv:2604.16909v1 Announce Type: new Abstract: As large language models (LLMs) evolve from conversational assistants into agents capable of handling complex tasks, they are increasingly deployed in h

model-releasesarxiv-cs-cl
21 Apr 2026
Research

PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues

DGX agent

arXiv:2604.18354v1 Announce Type: new Abstract: Emotion plays a pivotal role in shaping negotiation outcomes, influencing trust, cooperation, and long-term relationships. Developing negotiation dialog

researcharxiv-cs-cl
21 Apr 2026
Safety

Privacy Collapse: Benign Fine-Tuning Can Break Contextual Privacy in Language Models

DGX agent

arXiv:2601.15220v2 Announce Type: replace Abstract: We identify a novel phenomenon in language models: benign fine-tuning of frontier models can lead to privacy collapse. We find that diverse, subtle

safetyarxiv-cs-cl
21 Apr 2026
Local Ai

Privacy-R1: Privacy-Aware Multi-LLM Agent Collaboration via Reinforcement Learning

DGX agent

arXiv:2510.16054v2 Announce Type: replace-cross Abstract: When users submit queries to Large Language Models (LLMs), their prompts can often contain sensitive data, forcing a difficult choice: Send th

local-aiarxiv-cs-cl
21 Apr 2026
Research

PRL: Prompts from Reinforcement Learning

DGX agent

arXiv:2505.14412v2 Announce Type: replace-cross Abstract: Effective prompt engineering remains a central challenge in fully harnessing the capabilities of LLMs. While well-designed prompts can dramati

researcharxiv-cs-cl
21 Apr 2026
Hardware

Probabilistic Programs of Thought

DGX agent

arXiv:2604.17290v1 Announce Type: new Abstract: LLMs are widely used for code generation and mathematical reasoning tasks where they are required to generate structured output. They either need to rea

hardwarearxiv-cs-cl
21 Apr 2026
Tutorials

Procedural Knowledge at Scale Improves Reasoning

DGX agent

arXiv:2604.01348v2 Announce Type: replace Abstract: Test-time scaling has emerged as an effective way to improve language models on challenging reasoning tasks. However, most existing methods treat ea

tutorialsarxiv-cs-cl
21 Apr 2026
Research

Process Reward Models Meet Planning: Generating Precise and Scalable Datasets for Step-Level Rewards

DGX agent

arXiv:2604.17957v1 Announce Type: new Abstract: Process Reward Models (PRMs) have emerged as a powerful tool for providing step-level feedback when evaluating the reasoning of Large Language Models (L

researcharxiv-cs-cl
21 Apr 2026
Model Releases

ProfVLM: A lightweight video-language model for multi-view proficiency estimation

DGX agent

arXiv:2509.26278v4 Announce Type: replace-cross Abstract: Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically pr

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Prune, Interpret, Evaluate: A Cross-Layer Transcoder-Native Framework for Efficient Circuit Discovery via Feature Attribution

DGX agent

arXiv:2604.16889v1 Announce Type: new Abstract: Existing feature-interpretation pipelines typically operate on uniformly sampled units, but only a small fraction of cross-layer transcoder (CLT) featur

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Pseudo2Real: Task Arithmetic for Pseudo-Label Correction in Automatic Speech Recognition

DGX agent

arXiv:2510.08047v2 Announce Type: replace-cross Abstract: Robust ASR under domain shift is crucial because real-world systems encounter unseen accents and domains with limited labeled data. Although p

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

QU-NLP at QIAS 2026: Multi-Stage QLoRA Fine-Tuning for Arabic Islamic Inheritance Reasoning

DGX agent

arXiv:2604.16396v1 Announce Type: new Abstract: Islamic inheritance law (ilm al-mawar{i}th) presents a challenging domain for evaluating large language models' structured reasoning capabilities, requi

model-releasesarxiv-cs-cl
21 Apr 2026
Research

QuickScope: Certifying Hard Questions in Dynamic LLM Benchmarks

DGX agent

arXiv:2604.17842v1 Announce Type: new Abstract: LLM benchmarks are increasingly dynamic: instead of containing a fixed set of questions, they define templates and parameters that can generate an effec

researcharxiv-cs-cl
21 Apr 2026
Applications

RA-RRG: Multimodal Retrieval-Augmented Radiology Report Generation with Key Phrase Extraction

DGX agent

arXiv:2504.07415v2 Announce Type: replace-cross Abstract: Automated radiology report generation (RRG) holds potential to reduce the workload of radiologists, and recent advances in multimodal large la

applicationsarxiv-cs-cl
21 Apr 2026
Applications

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation

DGX agent

arXiv:2604.16310v1 Announce Type: cross Abstract: Evaluating Retrieval-Augmented Generation (RAG) systems using static multi-turn datasets fails to capture the dynamic nature of real-world dialogues.

applicationsarxiv-cs-cl
21 Apr 2026
Model Releases

ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval

DGX agent

arXiv:2510.08252v2 Announce Type: replace-cross Abstract: In this paper, we introduce ReasonEmbed, a novel text embedding model developed for reasoning-intensive document retrieval. Our work includes

model-releasesarxiv-cs-cl
21 Apr 2026
Research

Reasoning Models Know What's Important, and Encode It in Their Activations

DGX agent

arXiv:2604.18307v1 Announce Type: new Abstract: Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are cr

researcharxiv-cs-cl
21 Apr 2026
Tutorials

Reciprocal Co-Training (RCT): Coupling Gradient-Based and Non-Differentiable Models via Reinforcement Learning

DGX agent

arXiv:2604.16378v1 Announce Type: new Abstract: Large language models (LLMs) and classical machine learning methods offer complementary strengths for predictive modeling, yet their fundamentally diffe

tutorialsarxiv-cs-cl
21 Apr 2026
Model Releases

ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering

DGX agent

arXiv:2604.17944v1 Announce Type: new Abstract: Developing agents capable of navigating fragmented, multi-source information remains challenging, primarily due to the scarcity of benchmarks reflecting

model-releasesarxiv-cs-cl
21 Apr 2026
Applications

REFLEX: Reference-Free Evaluation of Log Summarization via Large Language Model Judgment

DGX agent

arXiv:2511.07458v2 Announce Type: replace Abstract: Evaluating log summarization systems is challenging due to the lack of high-quality reference summaries and the limitations of existing metrics like

applicationsarxiv-cs-cl
21 Apr 2026
Model Releases

REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control

DGX agent

arXiv:2511.20233v3 Announce Type: replace Abstract: The prevalence of fake news on social media demands automated fact-checking systems to provide accurate verdicts with faithful explanations. However

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReflexiCoder: Teaching Large Language Models to Self-Reflect on Generated Code and Self-Correct It via Reinforcement Learning

DGX agent

arXiv:2603.05863v2 Announce Type: replace Abstract: While Large Language Models (LLMs) have revolutionized code generation, standard ``System 1'' approaches that generate solutions in a single forward

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Reinforced Efficient Reasoning via Semantically Diverse Exploration

DGX agent

arXiv:2601.05053v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has proven effective in enhancing the reasoning of large language models (LLMs). Monte C

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

Reliability-Aware Adaptive Self-Consistency for Efficient Sampling in LLM Reasoning

DGX agent

arXiv:2601.02970v2 Announce Type: replace Abstract: Self-Consistency improves reasoning reliability through multi-sample aggregation, but incurs substantial inference cost. Adaptive self-consistency m

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning

DGX agent

arXiv:2601.14750v3 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has achieved remarkable success in unlocking the reasoning capabilities of Large Language Models (LLMs). Although C

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Representation-Guided Parameter-Efficient LLM Unlearning

DGX agent

arXiv:2604.17396v1 Announce Type: new Abstract: Large Language Models (LLMs) often memorize sensitive or harmful information, necessitating effective machine unlearning techniques. While existing para

model-releasesarxiv-cs-cl
21 Apr 2026
Tutorials

RePrompT: Recurrent Prompt Tuning for Integrating Structured EHR Encoders with Large Language Models

DGX agent

arXiv:2604.17725v1 Announce Type: new Abstract: Large Language Models (LLMs) have shown strong promise for mining Electronic Health Records (EHRs) by reasoning over longitudinal clinical information t

tutorialsarxiv-cs-cl
21 Apr 2026
Model Releases

ResearchBench: Benchmarking LLMs in Scientific Discovery via Inspiration-Based Task Decomposition

DGX agent

arXiv:2503.21248v3 Announce Type: replace Abstract: Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses r

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Rethinking Jailbreak Detection of Large Vision Language Models with Representational Contrastive Scoring

DGX agent

arXiv:2512.12069v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) are vulnerable to a growing array of multimodal jailbreak attacks, necessitating defenses that are both g

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Rethinking Meeting Effectiveness: A Benchmark and Framework for Temporal Fine-grained Automatic Meeting Effectiveness Evaluation

DGX agent

arXiv:2604.17260v1 Announce Type: new Abstract: Evaluating meeting effectiveness is crucial for improving organizational productivity. Current approaches rely on post-hoc surveys that yield a single c

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering

DGX agent

arXiv:2510.09351v2 Announce Type: replace Abstract: While Small Language Models (SLMs) have demonstrated promising performance on an increasingly wide array of commonsense reasoning benchmarks, curren

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Retrieval-Augmented Multimodal Model for Fake News Detection

DGX agent

arXiv:2604.18112v1 Announce Type: new Abstract: In recent years, multimodal multidomain fake news detection has garnered increasing attention. Nevertheless, this direction presents two significant cha

safetyarxiv-cs-cl
21 Apr 2026
Safety

Reverse Constitutional AI: A Framework for Controllable Toxic Data Generation via Probability-Clamped RLAIF

DGX agent

arXiv:2604.17769v1 Announce Type: new Abstract: Ensuring the safety of large language models (LLMs) requires robust red teaming, yet the systematic synthesis of high-quality toxic data remains under-e

safetyarxiv-cs-cl
21 Apr 2026
Model Releases

Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models

DGX agent

arXiv:2604.16593v1 Announce Type: new Abstract: We present SemanticQA, an evaluation suite designed to assess language models (LMs) in semantic phrase processing tasks. The benchmark consolidates exis

model-releasesarxiv-cs-cl
21 Apr 2026
Local Ai

Revisiting Entropy in Reinforcement Learning for Large Reasoning Models

DGX agent

arXiv:2511.05993v3 Announce Type: replace Abstract: Reinforcement learning with verifiable rewards (RLVR) has emerged as a prominent paradigm for enhancing the reasoning capabilities of large language

local-aiarxiv-cs-cl
21 Apr 2026
Safety

REZE: Representation Regularization for Domain-adaptive Text Embedding Pre-finetuning

DGX agent

arXiv:2604.17257v1 Announce Type: new Abstract: Recent text embedding models are often adapted to specialized domains via contrastive pre-finetuning (PFT) on a naive collection of scattered, heterogen

safetyarxiv-cs-cl
21 Apr 2026
Tutorials

River-LLM: Large Language Model Seamless Exit Based on KV Share

DGX agent

arXiv:2604.18396v1 Announce Type: new Abstract: Large Language Models (LLMs) have demonstrated exceptional performance across diverse domains but are increasingly constrained by high inference latency

tutorialsarxiv-cs-cl
21 Apr 2026
Model Releases

Robust Bias Evaluation with FilBBQ: A Filipino Bias Benchmark for Question-Answering Language Models

DGX agent

arXiv:2602.14466v2 Announce Type: replace Abstract: With natural language generation becoming a popular use case for language models, the Bias Benchmark for Question-Answering (BBQ) has grown to be an

model-releasesarxiv-cs-cl
21 Apr 2026
Model Releases

RoIt-XMASA: Multi-Domain Multilingual Sentiment Analysis Dataset for Romanian and Italian

DGX agent

arXiv:2604.17134v1 Announce Type: new Abstract: We present RoIt-XMASA, a multilingual dataset that extends the Cross-lingual Multi-domain Amazon Sentiment Analysis to Italian and Romanian, comprising

model-releasesarxiv-cs-cl
21 Apr 2026
← Previous
1…136137138139140…160
Next →