AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Research

Behavior-Aware Item Modeling via Dynamic Procedural Solution Representations for Knowledge Tracing

DGX agent

arXiv:2604.08260v1 Announce Type: new Abstract: Knowledge Tracing (KT) aims to predict learners' future performance from past interactions. While recent KT approaches have improved via learning item r

researcharxiv-cs-cl
10 Apr 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

BenchBrowser: Retrieving Evidence for Evaluating Benchmark Validity

DGX agent

arXiv:2603.18019v2 Announce Type: replace Abstract: Do language model benchmarks actually measure what practitioners intend them to ? High-level metadata is too coarse to convey the granular reality o

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Beyond Social Pressure: Benchmarking Epistemic Attack in Large Language Models

DGX agent

arXiv:2604.07749v1 Announce Type: new Abstract: Large language models (LLMs) can shift their answers under pressure in ways that reflect accommodation rather than reasoning. Prior work on sycophancy h

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Break Me If You Can: Self-Jailbreaking of Aligned LLMs via Lexical Insertion Prompting

DGX agent

arXiv:2601.02670v2 Announce Type: replace Abstract: We introduce self-jailbreaking, a threat model in which an aligned LLM guides its own compromise. Unlike most jailbreak techniques, which oft

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

CAMO: A Class-Aware Minority-Optimized Ensemble for Robust Language Model Evaluation on Imbalanced Data

DGX agent

arXiv:2604.07583v1 Announce Type: new Abstract: Real-world categorization is severely hampered by class imbalance because traditional ensembles favor majority classes, which lowers minority performanc

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Can Vision Language Models Judge Action Quality? An Empirical Evaluation

DGX agent

arXiv:2604.08294v1 Announce Type: cross Abstract: Action Quality Assessment (AQA) has broad applications in physical therapy, sports coaching, and competitive judging. Although Vision Language Models

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Chunks as Arms: Multi-Armed Bandit-Guided Sampling for Long-Context LLM Preference Optimization

DGX agent

arXiv:2508.13993v2 Announce Type: replace Abstract: Long-context modeling is critical for a wide range of real-world tasks, including long-context question answering, summarization, and complex reason

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

ClawBench: Can AI Agents Complete Everyday Online Tasks?

DGX agent

arXiv:2604.08523v1 Announce Type: new Abstract: AI agents may be able to automate your inbox, but can they automate other routine aspects of your life? Everyday online tasks offer a realistic yet unso

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Clickbait detection: quick inference with maximum impact

DGX agent

arXiv:2604.08148v1 Announce Type: new Abstract: We propose a lightweight hybrid approach to clickbait detection that combines OpenAI semantic embeddings with six compact heuristic features capturing s

researcharxiv-cs-cl
10 Apr 2026
Local Ai

Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs

DGX agent

arXiv:2603.20698v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated remarkable potential in medical image analysis. However, their application in gastr

local-aiarxiv-cs-cl
10 Apr 2026
Research

Compact Example-Based Explanations for Language Models

DGX agent

arXiv:2601.03786v2 Announce Type: replace Abstract: Training data influence estimation methods quantify the contribution of training documents to a model's output, making them a promising source of in

researcharxiv-cs-cl
10 Apr 2026
Model Releases

ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training

DGX agent

arXiv:2604.07484v1 Announce Type: cross Abstract: Generative reward models (GRMs) have emerged as a promising approach for aligning Large Language Models (LLMs) with human preferences by offering grea

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Contextual Earnings-22: A Speech Recognition Benchmark with Custom Vocabulary in the Wild

DGX agent

arXiv:2604.07354v1 Announce Type: new Abstract: The accuracy frontier of speech-to-text systems has plateaued on academic benchmarks.1 In contrast, industrial benchmarks and adoption in high-stakes do

model-releasesarxiv-cs-cl
10 Apr 2026
Safety

Contextualising (Im)plausible Events Triggers Figurative Language

DGX agent

arXiv:2604.07885v1 Announce Type: new Abstract: This work explores the connection between (non-)literalness and plausibility at the example of subject-verb-object events in English. We design a system

safetyarxiv-cs-cl
10 Apr 2026
Research

Cram Less to Fit More: Training Data Pruning Improves Memorization of Facts

DGX agent

arXiv:2604.08519v1 Announce Type: new Abstract: Large language models (LLMs) can struggle to memorize factual knowledge in their parameters, often leading to hallucinations and poor performance on kno

researcharxiv-cs-cl
10 Apr 2026
Research

Cross-Tokenizer LLM Distillation through a Byte-Level Interface

DGX agent

arXiv:2604.07466v1 Announce Type: new Abstract: Cross-tokenizer distillation (CTD), the transfer of knowledge from a teacher to a student language model when the two use different tokenizers, remains

researcharxiv-cs-cl
10 Apr 2026
Research

Current LLMs still cannot 'talk much' about grammar modules: Evidence from syntax

DGX agent

arXiv:2603.20114v4 Announce Type: replace Abstract: We aim to examine the extent to which Large Language Models (LLMs) can 'talk much' about grammar modules, providing evidence from syntax core proper

researcharxiv-cs-cl
10 Apr 2026
Model Releases

CycleChart: A Unified Consistency-Based Learning Framework for Bidirectional Chart Understanding and Generation

DGX agent

arXiv:2512.19173v2 Announce Type: replace Abstract: Current chart-related tasks, such as chart generation (NL2Chart), chart schema parsing, chart data parsing, and chart question answering (ChartQA),

model-releasesarxiv-cs-cl
10 Apr 2026
Local Ai

Data Selection for Multi-turn Dialogue Instruction Tuning

DGX agent

arXiv:2604.07892v1 Announce Type: new Abstract: Instruction-tuned language models increasingly rely on large multi-turn dialogue corpora, but these datasets are often noisy and structurally inconsiste

local-aiarxiv-cs-cl
10 Apr 2026
Safety

Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs

DGX agent

arXiv:2604.07518v1 Announce Type: new Abstract: Vision-Language Models often struggle with complex visual reasoning due to the visual information loss in textual CoT. Existing methods either add the c

safetyarxiv-cs-cl
10 Apr 2026
Safety

Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models

DGX agent

arXiv:2604.08527v1 Announce Type: new Abstract: On-policy distillation (OPD) trains student models under their own induced distribution while leveraging supervision from stronger teachers. We identify

safetyarxiv-cs-cl
10 Apr 2026
Model Releases

Detecting HIV-Related Stigma in Clinical Narratives Using Large Language Models

DGX agent

arXiv:2604.07717v1 Announce Type: new Abstract: Human immunodeficiency virus (HIV)-related stigma is a critical psychosocial determinant of health for people living with HIV (PLWH), influencing mental

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Differentially Private Language Generation and Identification in the Limit

DGX agent

arXiv:2604.08504v1 Announce Type: cross Abstract: We initiate the study of language generation in the limit, a model recently introduced by Kleinberg and Mullainathan [KM24], under the constraint of d

researcharxiv-cs-cl
10 Apr 2026
Research

Diffusion Language Models Know the Answer Before Decoding

DGX agent

arXiv:2508.19982v5 Announce Type: replace Abstract: Diffusion language models (DLMs) have recently emerged as an alternative to autoregressive approaches, offering parallel sequence generation and fle

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Distributed Multi-Layer Editing for Rule-Level Knowledge in Large Language Models

DGX agent

arXiv:2604.08284v1 Announce Type: new Abstract: Large language models store not only isolated facts but also rules that support reasoning across symbolic expressions, natural language explanations, an

model-releasesarxiv-cs-cl
10 Apr 2026
Research

DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification

DGX agent

arXiv:2604.07622v1 Announce Type: new Abstract: Speculative decoding is an effective technique for accelerating large language model inference by drafting multiple tokens in parallel. In practice, its

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Don't Overthink It: Inter-Rollout Action Agreement as a Free Adaptive-Compute Signal for LLM Agents

DGX agent

arXiv:2604.08369v1 Announce Type: cross Abstract: Inference-time compute scaling has emerged as a powerful technique for improving the reliability of large language model (LLM) agents, but existing me

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism

DGX agent

arXiv:2601.05524v2 Announce Type: replace Abstract: Parallel Speculative Decoding (PSD) accelerates traditional Speculative Decoding (SD) by overlapping draft generation with verification. However, it

researcharxiv-cs-cl
10 Apr 2026
Applications

DQA: Diagnostic Question Answering for IT Support

DGX agent

arXiv:2604.05350v2 Announce Type: replace Abstract: Enterprise IT support interactions are fundamentally diagnostic: effective resolution requires iterative evidence gathering from ambiguous user repo

applicationsarxiv-cs-cl
10 Apr 2026
Model Releases

Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving

DGX agent

arXiv:2604.08075v1 Announce Type: new Abstract: Production vLLM fleets typically provision each instance for the worst-case context length, leading to substantial KV-cache over-allocation and under-ut

model-releasesarxiv-cs-cl
10 Apr 2026
Research

DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs

DGX agent

arXiv:2601.07994v4 Announce Type: replace Abstract: Large Language Models (LLMs) increasingly operate over long-form dialogues with frequent topic shifts. While recent LLMs support extended context wi

researcharxiv-cs-cl
10 Apr 2026
Model Releases

E2Edev: Benchmarking Large Language Models in End-to-End Software Development Task

DGX agent

arXiv:2510.14509v3 Announce Type: replace-cross Abstract: The rapid advancement in large language models (LLMs) has demonstrated significant potential in End-to-End Software Development (E2ESD). Howev

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction

DGX agent

arXiv:2604.07659v1 Announce Type: new Abstract: Large language models (LLMs) hold significant promise for healthcare, yet their reliability in high-stakes clinical settings is often compromised by hal

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Efficient PRM Training Data Synthesis via Formal Verification

DGX agent

arXiv:2505.15960v3 Announce Type: replace Abstract: Process Reward Models (PRMs) have emerged as a promising approach for improving LLM reasoning capabilities by providing process supervision over rea

researcharxiv-cs-cl
10 Apr 2026
Research

Efficient Provably Secure Linguistic Steganography via Range Coding

DGX agent

arXiv:2604.08052v1 Announce Type: new Abstract: Linguistic steganography involves embedding secret messages within seemingly innocuous texts to enable covert communication. Provable security, which is

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Emotion Concepts and their Function in a Large Language Model

DGX agent

arXiv:2604.07729v1 Announce Type: cross Abstract: Large language models (LLMs) sometimes appear to exhibit emotional reactions. We investigate why this is the case in Claude Sonnet 4.5 and explore imp

model-releasesarxiv-cs-cl
10 Apr 2026
Agents

EMSDialog: Synthetic Multi-person Emergency Medical Service Dialogue Generation from Electronic Patient Care Reports via Multi-LLM Agents

DGX agent

arXiv:2604.07549v1 Announce Type: new Abstract: Conversational diagnosis prediction requires models to track evolving evidence in streaming clinical conversations and decide when to commit to a diagno

agentsarxiv-cs-cl
10 Apr 2026
Model Releases

Enabling Intrinsic Reasoning over Dense Geospatial Embeddings with DFR-Gemma

DGX agent

arXiv:2604.07490v1 Announce Type: new Abstract: Representation learning for geospatial and spatio-temporal data plays a critical role in enabling general-purpose geospatial intelligence. Recent geospa

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models

DGX agent

arXiv:2604.08456v1 Announce Type: cross Abstract: Despite rapid progress, pretrained vision-language models still struggle when answers depend on tiny visual details or on combining clues spread acros

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Evaluating LLMs for Demographic-Targeted Social Bias Detection: A Comprehensive Benchmark Study

DGX agent

arXiv:2510.04641v3 Announce Type: replace Abstract: Large-scale web-scraped text corpora used to train general-purpose AI models often contain harmful demographic-targeted social biases, creating a re

model-releasesarxiv-cs-cl
10 Apr 2026
Tutorials

EventWeave: A Dynamic Framework for Capturing Core and Supporting Events in Dialogue Systems

DGX agent

arXiv:2503.23078v3 Announce Type: replace Abstract: Large language models have improved dialogue systems, but often process conversational turns in isolation, overlooking the event structures that gui

tutorialsarxiv-cs-cl
10 Apr 2026
Agents

Exploring Plan Space through Conversation: An Agentic Framework for LLM-Mediated Explanations in Planning

DGX agent

arXiv:2603.02070v2 Announce Type: replace-cross Abstract: When automating plan generation for a real-world sequential decision problem, the goal is often not to replace the human planner, but to facil

agentsarxiv-cs-cl
10 Apr 2026
Model Releases

FinTruthQA: A Benchmark for AI-Driven Financial Disclosure Quality Assessment in Investor -- Firm Interactions

DGX agent

arXiv:2406.12009v5 Announce Type: replace Abstract: Accurate and transparent financial information disclosure is essential for market efficiency, investor decision-making, and corporate governance. Ch

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Floating or Suggesting Ideas? A Large-Scale Contrastive Analysis of Metaphorical and Literal Verb-Object Constructions

DGX agent

arXiv:2604.08275v1 Announce Type: new Abstract: Metaphor pervades everyday language, allowing speakers to express abstract concepts via concrete domains. While prior work has studied metaphors cogniti

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference

DGX agent

arXiv:2604.07394v1 Announce Type: cross Abstract: The quadratic computational complexity of standard attention mechanisms presents a severe scalability bottleneck for LLMs in long-context scenarios. W

model-releasesarxiv-cs-cl
10 Apr 2026
Research

Formalizing building-up constructions of self-dual codes through isotropic lines in Lean

DGX agent

arXiv:2604.08485v1 Announce Type: cross Abstract: The purpose of this paper is two-fold. First we show that Kim's building-up construction of binary self-dual codes is equivalent to Chinburg-Zhang's H

researcharxiv-cs-cl
10 Apr 2026
Model Releases

From Fragments to Facts: A Curriculum-Driven DPO Approach for Generating Hindi News Veracity Explanations

DGX agent

arXiv:2507.05179v4 Announce Type: replace Abstract: In an era of rampant misinformation, generating reliable news explanations is vital, especially for under-represented languages like Hindi. Lacking

model-releasesarxiv-cs-cl
10 Apr 2026
Safety

From Ground Truth to Measurement: A Statistical Framework for Human Labeling

DGX agent

arXiv:2604.07591v1 Announce Type: cross Abstract: Supervised machine learning assumes that labeled data provide accurate measurements of the concepts models are meant to learn. Yet in practice, human

safetyarxiv-cs-cl
10 Apr 2026
← Previous
1…155156157158159160
Next →