AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
17 Apr 2026

CobwebTM: Probabilistic Concept Formation for Lifelong and Hierarchical Topic Modeling

Model ReleasesDGX agent

arXiv:2604.14489v1 Announce Type: new Abstract: Topic modeling seeks to uncover latent semantic structure in text corpora with minimal supervision. Neural approaches achieve strong performance but req

Cognitive Alpha Mining via LLM-Driven Code-Based Evolution

ResearchDGX agent

arXiv:2511.18850v2 Announce Type: replace Abstract: Discovering effective predictive signals, or 'alphas,' from financial data with high dimensionality and extremely low signal-to-noise ratio remains

Comparison of Modern Multilingual Text Embedding Techniques for Hate Speech Detection Task

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2604.14907v1 Announce Type: new Abstract: Online hate speech and abusive language pose a growing challenge for content moderation, especially in multilingual settings and for low-resource langua

Compressed-Sensing-Guided, Inference-Aware Structured Reduction for Large Language Models

Model ReleasesDGX agent

arXiv:2604.14156v1 Announce Type: new Abstract: Large language models deliver strong generative performance but at the cost of massive parameter counts, memory use, and decoding latency. Prior work ha

Compressing Sequences in the Latent Embedding Space: K-Token Merging for Large Language Models

ResearchDGX agent

arXiv:2604.15153v1 Announce Type: new Abstract: Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically

ConfLayers: Adaptive Confidence-based Layer Skipping for Self-Speculative Decoding

SafetyDGX agent

arXiv:2604.14612v1 Announce Type: cross Abstract: Self-speculative decoding is an inference technique for large language models designed to speed up generation without sacrificing output quality. It c

Context Over Content: Exposing Evaluation Faking in Automated Judges

SafetyDGX agent

arXiv:2604.15224v1 Announce Type: cross Abstract: The extit{LLM-as-a-judge} paradigm has become the operational backbone of automated AI evaluation pipelines, yet rests on an unverified assumption: th

Controlling Authority Retrieval: A Missing Retrieval Objective for Authority-Governed Knowledge

Model ReleasesDGX agent

arXiv:2604.14488v1 Announce Type: cross Abstract: In any domain where knowledge accumulates under formal authority -- law, drug regulation, software security -- a later document can formally void an e

CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas

SafetyDGX agent

arXiv:2604.15267v1 Announce Type: cross Abstract: It is increasingly important that LLM agents interact effectively and safely with other goal-pursuing agents, yet, recent works report the opposite tr

CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors

Model ReleasesDGX agent

arXiv:2604.14773v1 Announce Type: new Abstract: While LLMs have demonstrated remarkable potential in Question Answering (QA), evaluating personalization remains a critical bottleneck. Existing paradig

Correcting Suppressed Log-Probabilities in Language Models with Post-Transformer Adapters

Model ReleasesDGX agent

arXiv:2604.14174v1 Announce Type: new Abstract: Alignment-tuned language models frequently suppress factual log-probabilities on politically sensitive topics despite retaining the knowledge in their h

Cosine-Similarity Routing with Semantic Anchors for Interpretable Mixture-of-Experts Language Models

ResearchDGX agent

arXiv:2509.14255v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) models improve efficiency through sparse activation, but their learned gating functions provide limited insight into routin

Counting Without Numbers and Finding Without Words

ResearchDGX agent

arXiv:2603.24470v2 Announce Type: replace-cross Abstract: Every year, 10 million pets enter shelters, separated from their families. Despite desperate searches by both guardians and lost animals, 70%

CROP: Token-Efficient Reasoning in Large Language Models via Regularized Prompt Optimization

AgentsDGX agent

arXiv:2604.14214v1 Announce Type: new Abstract: Large Language Models utilizing reasoning techniques improve task performance but incur significant latency and token costs due to verbose generation. E

CURA: Clinical Uncertainty Risk Alignment for Language Model-Based Risk Prediction

SafetyDGX agent

arXiv:2604.14651v1 Announce Type: new Abstract: Clinical language models (LMs) are increasingly applied to support clinical risk prediction from free-text notes, yet their uncertainty estimates often

CURaTE: Continual Unlearning in Real Time with Ensured Preservation of LLM Knowledge

ResearchDGX agent

arXiv:2604.14644v1 Announce Type: new Abstract: The inability to filter out in advance all potentially problematic data from the pre-training of large language models has given rise to the need for me

DA-Cramming: Enhancing Cost-Effective Language Model Pretraining with Dependency Agreement Integration

HardwareDGX agent

arXiv:2311.04799v2 Announce Type: replace Abstract: Pretraining language models is still a challenge for many researchers due to its substantial computational costs. As such, there is growing interest

Dark & Stormy: Modeling Humor in Sentences from the Bulwer-Lytton Fiction Contest

ResearchDGX agent

arXiv:2510.24538v2 Announce Type: replace Abstract: Textual humor is enormously diverse and computational studies need to account for this range, including intentionally bad humor. In this paper, we c

De-Anonymization at Scale via Tournament-Style Attribution

ApplicationsDGX agent

arXiv:2601.12407v2 Announce Type: replace-cross Abstract: As LLMs rapidly advance and enter real-world use, their privacy implications are increasingly important. We study an authorship de-anonymizati

Decoupling Scores and Text: The Politeness Principle in Peer Review

ResearchDGX agent

arXiv:2604.14162v1 Announce Type: new Abstract: Authors often struggle to interpret peer review feedback, deriving false hope from polite comments or feeling confused by specific low scores. To invest

DeepPrune: Parallel Scaling without Inter-trace Redundancy

ResearchDGX agent

arXiv:2510.08483v2 Announce Type: replace Abstract: Parallel scaling has emerged as a powerful paradigm to enhance reasoning capabilities in large language models (LLMs) by generating multiple Chain-o

DharmaOCR: Specialized Small Language Models for Structured OCR that outperform Open-Source and Commercial Baselines

Model ReleasesDGX agent

arXiv:2604.14314v1 Announce Type: cross Abstract: This manuscript introduces DharmaOCR Full and Lite, a pair of specialized small language models (SSLMs) for structured OCR that jointly optimize trans

Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations

ResearchDGX agent

arXiv:2604.15302v1 Announce Type: cross Abstract: LLM-as-judge frameworks are increasingly used for automatic NLG evaluation, yet their per-instance reliability remains poorly understood. We present a

DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering

TutorialsDGX agent

arXiv:2604.15140v1 Announce Type: new Abstract: We introduce DiscoTrace, a method to identify the rhetorical strategies that answerers use when responding to information-seeking questions. DiscoTrace

Dissecting Failure Dynamics in Large Language Model Reasoning

Local AiDGX agent

arXiv:2604.14528v1 Announce Type: cross Abstract: Large Language Models (LLMs) achieve strong performance through extended inference-time deliberation, yet how their reasoning failures arise remains p

Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems

Model ReleasesDGX agent

arXiv:2604.14228v1 Announce Type: cross Abstract: Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes

Domain Fine-Tuning FinBERT on Finnish Histopathological Reports: Train-Time Signals and Downstream Correlations

ApplicationsDGX agent

arXiv:2604.14815v1 Announce Type: new Abstract: In NLP classification tasks where little labeled data exists, domain fine-tuning of transformer models on unlabeled data is an established approach. In

Don't Retrieve, Navigate: Distilling Enterprise Knowledge into Navigable Agent Skills for QA and RAG

Model ReleasesDGX agent

arXiv:2604.14572v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) grounds LLM responses in external evidence but treats the model as a passive consumer of search results: it never

DySCO: Dynamic Attention-Scaling Decoding for Long-Context Language Models

ResearchDGX agent

arXiv:2602.22175v2 Announce Type: replace Abstract: Understanding and reasoning over long contexts is a crucial capability for language models (LMs). Although recent models support increasingly long c

Enhancing Linguistic Competence of Language Models through Pre-training with Language Learning Tasks

ResearchDGX agent

arXiv:2601.03448v2 Announce Type: replace Abstract: Language models (LMs) are pre-trained on raw text datasets to generate text sequences token-by-token. While this approach facilitates the learning o

EuropeMedQA Study Protocol: A Multilingual, Multimodal Medical Examination Dataset for Language Model Evaluation

Model ReleasesDGX agent

arXiv:2604.14306v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated high proficiency on English-centric medical examinations, their performance often declines when fac

EviSearch: A Human in the Loop System for Extracting and Auditing Clinical Evidence for Systematic Reviews

Model ReleasesDGX agent

arXiv:2604.14165v1 Announce Type: new Abstract: We present EviSearch, a multi-agent extraction system that automates the creation of ontology-aligned clinical evidence tables directly from native tria

Evolving Beyond Snapshots: Harmonizing Structure and Sequence via Entity State Tuning for Temporal Knowledge Graph Forecasting

ResearchDGX agent

arXiv:2602.12389v3 Announce Type: replace-cross Abstract: Temporal knowledge graph (TKG) forecasting requires predicting future facts by jointly modeling structural dependencies within each snapshot a

Explain the Flag: Contextualizing Hate Speech Beyond Censorship

ResearchDGX agent

arXiv:2604.14970v1 Announce Type: new Abstract: Hate, derogatory, and offensive speech remains a persistent challenge in online platforms and public discourse. While automated detection systems are wi

Exploring and Testing Skill-Based Behavioral Profile Annotation: Human Operability and LLM Feasibility under Schema-Guided Execution

Model ReleasesDGX agent

arXiv:2604.14843v1 Announce Type: new Abstract: Behavioral Profile (BP) annotation is difficult to automate because it requires simultaneous coding across multiple linguistic dimensions. We treat BP a

Fabricator or dynamic translator?

ResearchDGX agent

arXiv:2604.15165v1 Announce Type: new Abstract: LLMs are proving to be adept at machine translation although due to their generative nature they may at times overgenerate in various ways. These overge

Fact4ac at the Financial Misinformation Detection Challenge Task: Reference-Free Financial Misinformation Detection via Fine-Tuning and Few-Shot Prompting of Large Language Models

Model ReleasesDGX agent

arXiv:2604.14640v1 Announce Type: new Abstract: The proliferation of financial misinformation poses a severe threat to market stability and investor trust, misleading market behavior and creating crit

Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance

ResearchDGX agent

arXiv:2604.14325v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong performance and have revolutionized NLP, but their lack of explainability keeps them treated as black boxes,

Feedback Adaptation for Retrieval-Augmented Generation

ResearchDGX agent

arXiv:2604.06647v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) systems are typically evaluated under static assumptions, despite being frequently corrected through user or ex

Filling in the Mechanisms: How do LMs Learn Filler-Gap Dependencies under Developmental Constraints?

SafetyDGX agent

arXiv:2604.14459v1 Announce Type: new Abstract: For humans, filler-gap dependencies require a shared representation across different syntactic constructions. Although causal analyses suggest this may

From Black Box to Glass Box: Cross-Model ASR Disagreement to Prioto Review in Ambient AI Scribe Documentation

ApplicationsDGX agent

arXiv:2604.14152v1 Announce Type: cross Abstract: Ambient AI 'scribe' systems promise to reduce clinical documentation burden, but automatic speech recognition (ASR) errors can remain unnoticed withou

From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities

SafetyDGX agent

arXiv:2604.03920v2 Announce Type: replace Abstract: LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tes

From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution

TutorialsDGX agent

arXiv:2604.15097v1 Announce Type: cross Abstract: This beta technical report asks how reusable experience should be represented so that it can function as effective test-time control and as a substrat

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench

ResearchDGX agent

arXiv:2604.15037v1 Announce Type: cross Abstract: Recent advancements in LLM agents are gradually shifting from reactive, text-based paradigms toward proactive, multimodal interaction. However, existi

From Tokens to Steps: Verification-Aware Speculative Decoding for Efficient Multi-Step Reasoning

ResearchDGX agent

arXiv:2604.15244v1 Announce Type: new Abstract: Speculative decoding (SD) accelerates large language model inference by allowing a lightweight draft model to propose outputs that a stronger target mod

Generating Concept Lexicalizations via Dictionary-Based Cross-Lingual Sense Projection

ResearchDGX agent

arXiv:2604.14397v1 Announce Type: new Abstract: We study the task of automatically expanding WordNet-style lexical resources to new languages through sense generation. We generate senses by associatin

Grading the Unspoken: Evaluating Tacit Reasoning in Quantum Field Theory and String Theory with LLMs

ResearchDGX agent

arXiv:2604.14188v1 Announce Type: cross Abstract: Large language models have demonstrated impressive performance across many domains of mathematics and physics. One natural question is whether such mo

Graph-Based Alternatives to LLMs for Human Simulation

ResearchDGX agent

arXiv:2511.02135v2 Announce Type: replace Abstract: Large language models (LLMs) have become a popular approach for simulating human behaviors, yet it remains unclear if LLMs are necessary for all sim

HARNESS: Lightweight Distilled Arabic Speech Foundation Models

ApplicationsDGX agent

arXiv:2604.14186v1 Announce Type: cross Abstract: Large self-supervised speech (SSL) models achieve strong downstream performance, but their size limits deployment in resource-constrained settings. We

HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

HardwareDGX agent

arXiv:2601.14724v3 Announce Type: replace-cross Abstract: Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated significant improvement in offline video understanding. Howe

Hierarchical Retrieval Augmented Generation for Adversarial Technique Annotation in Cyber Threat Intelligence Text

SafetyDGX agent

arXiv:2604.14166v1 Announce Type: new Abstract: Mapping Cyber Threat Intelligence (CTI) text to MITRE ATT&CK technique IDs is a critical task for understanding adversary behaviors and automating threa

Hierarchical Semantic Retrieval with Cobweb

ResearchDGX agent

arXiv:2510.02539v2 Announce Type: replace Abstract: Neural document retrieval often treats a corpus as a flat cloud of vectors scored at a single granularity, leaving corpus structure underused and ex

Hierarchical vs. Flat Iteration in Shared-Weight Transformers

Model ReleasesDGX agent

arXiv:2604.14442v1 Announce Type: new Abstract: We present an empirical study of whether hierarchically structured, shared-weight recurrence can match the representational quality of independent-layer

How Retrieved Context Shapes Internal Representations in RAG

ResearchDGX agent

arXiv:2602.20091v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) enhances large language models (LLMs) by conditioning generation on retrieved external documents, but the effec

How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data

TutorialsDGX agent

arXiv:2604.14164v1 Announce Type: new Abstract: A widely adopted strategy for model enhancement is to use synthetic data generated by a stronger model for supervised fine-tuning (SFT). However, for em

HUOZIIME: An On-Device LLM-enhanced Input Method for Deep Personalization

Local AiDGX agent

arXiv:2604.14159v1 Announce Type: new Abstract: Mobile input method editors (IMEs) are the primary interface for text input, yet they remain constrained to manual typing and struggle to produce person

Hybrid Decision Making via Conformal VLM-generated Guidance

TutorialsDGX agent

arXiv:2604.14980v1 Announce Type: cross Abstract: Building on recent advances in AI, hybrid decision making (HDM) holds the promise of improving human decision quality and reducing cognitive load. We

IE as Cache: Information Extraction Enhanced Agentic Reasoning

AgentsDGX agent

arXiv:2604.14930v1 Announce Type: new Abstract: Information Extraction aims to distill structured, decision-relevant information from unstructured text, serving as a foundation for downstream understa

IF-CRITIC: Towards a Fine-Grained LLM Critic for Instruction-Following Evaluation

Model ReleasesDGX agent

arXiv:2511.01014v3 Announce Type: replace Abstract: Instruction-following is a fundamental ability of Large Language Models (LLMs), requiring their generated outputs to follow multiple constraints imp

IF-RewardBench: Benchmarking Judge Models for Instruction-Following Evaluation

Model ReleasesDGX agent

arXiv:2603.04738v2 Announce Type: replace Abstract: Instruction-following is a foundational capability of large language models (LLMs), with its improvement hinging on scalable and accurate feedback f

← Previous
1…113114115116117…128
Next →