AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
31 Jul 2026

Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention

ResearchDGX agent

arXiv:2607.27692v1 Announce Type: new Abstract: Top-K sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identi

ReDiPPO: Reference-Guided Value Calibration and Discrepancy-Aware Token Reweighting for Mathematical Reasoning

SafetyDGX agent

arXiv:2607.27631v1 Announce Type: cross Abstract: Reinforcement learning has emerged as an effective paradigm for enhancing the mathematical reasoning capabilities of large language models. Among exis

RepBench: Compiling Benchmarks into Capability Representations for Large Language Models

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2607.28008v1 Announce Type: new Abstract: Representation engineering reads and steers capability directions in large language models, yet methods are typically evaluated on paper-specific synthe

Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

Model ReleasesDGX agent

arXiv:2607.28128v1 Announce Type: new Abstract: LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit

RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning

ResearchDGX agent

arXiv:2607.28156v1 Announce Type: new Abstract: Existing multimodal long-term memory agents use external memory to overcome the limited context available for long videos. However, most methods emphasi

Safety Verification of Wait-Only Non-Blocking Broadcast Protocols

SafetyDGX agent

arXiv:2403.18591v3 Announce Type: replace-cross Abstract: Broadcast protocols are programs designed to be executed by networks of processes. Each process runs the same protocol, and communication betw

Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models

Model ReleasesDGX agent

arXiv:2607.27384v1 Announce Type: new Abstract: Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure

Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B

ResearchDGX agent

arXiv:2607.28576v1 Announce Type: new Abstract: Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with co

Scaling medical imaging report generation with multimodal reinforcement learning

Model ReleasesDGX agent

arXiv:2601.17151v2 Announce Type: replace-cross Abstract: Frontier models have demonstrated remarkable capabilities in understanding and reasoning with natural-language text, but they still exhibit ma

SciDataSailor: Deep Scientific Data Exploring

AgentsDGX agent

arXiv:2607.28098v1 Announce Type: cross Abstract: Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, in

SciSchema.org: A Multidisciplinary Collection of Schemas for Structured Scientific Process Descriptions

ResearchDGX agent

arXiv:2607.27955v1 Announce Type: cross Abstract: Scientific processes are often described in heterogeneous article discourse, with details needed for comparison, reproducibility, reuse, and automatio

Select or Project? Evaluating Lower-dimensional Vectors for LLM Training Data Explanations

Model ReleasesDGX agent

arXiv:2601.16651v3 Announce Type: replace Abstract: Gradient-based methods for instance-based explanation for large language models (LLMs) are hindered by the immense dimensionality of model gradients

Selecting Open-Weight Language Models for Zero-Shot Intent Classification: A Systematic Evaluation of 41 Models

Model ReleasesDGX agent

arXiv:2607.27421v1 Announce Type: new Abstract: Intent classification is a core component of task-oriented dialogue systems, yet practitioners have limited systematic guidance for selecting deployable

Semantic-Aligned Structural Abstraction for Multimodal Sentiment Analysis

ResearchDGX agent

arXiv:2607.27790v1 Announce Type: new Abstract: Multimodal Sentiment Analysis (MSA) aims to interpret complex human emotions by integrating natural language with non-verbal modalities. Non-verbal moda

SkillSmith: Learning to Compose Parametric Skills and Textual Knowledge

AgentsDGX agent

arXiv:2607.27497v1 Announce Type: new Abstract: Agentic systems driven by large language models (LLMs) regularly feature two key mechanisms to autonomously solve complex problems: synthesizing text-ba

Stage-Replay Divergence Follows the KV Cache: Fixed-Prefix Precision Controls and Bidirectional Cache Transplantation

ResearchDGX agent

arXiv:2607.28495v1 Announce Type: cross Abstract: Stage-replay diagnostics reconstruct intermediate token prefixes and treat fresh-prefill continuation as continuation from the decoder state that orig

Subtract or Replay? Exact Deletion from Language-Model Memory

Model ReleasesDGX agent

arXiv:2607.27539v1 Announce Type: cross Abstract: Exact deletion from persistent language-model memory depends on how that memory represents a record. Addressable influence can be removed by algebraic

SVR: Self-Verifying Refinement via Joint Verdict-Confidence Reinforcement Learning for Adaptive Test-Time Compute

SafetyDGX agent

arXiv:2607.28457v1 Announce Type: cross Abstract: Scaling test-time computation can improve language-model reasoning, but uniform budgets waste computation on easy inputs, while verifier-guided refine

Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups

Model ReleasesDGX agent

arXiv:2607.27232v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs

TCA-SIR: Learning Target-Conditioned Abstractions for Scientific Inspiration Retrieval

ResearchDGX agent

arXiv:2607.28498v1 Announce Type: cross Abstract: Scientific hypothesis generation for AI for Science typically involves Scientific Inspiration Retrieval (SIR) followed by hypothesis composition. Exis

The Confidence Manifold: Geometric Structure of Correctness Representations in Language Models

SafetyDGX agent

arXiv:2602.08159v2 Announce Type: replace-cross Abstract: When a language model asserts that 'the capital of Australia is Sydney,' does it know this is wrong? Models assert misconceptions with the sam

The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials

ResearchDGX agent

arXiv:2607.28190v1 Announce Type: new Abstract: Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the p

THGFM: Dual-Branch Temporal Heterogeneous Graph Fusion Model

Model ReleasesDGX agent

arXiv:2607.27303v1 Announce Type: cross Abstract: Temporal heterogeneous graphs offer a natural abstraction for dynamic relational systems in which diverse node and relation types co-exist and evolve

ThreatForest: Multi-Agent Attack Tree Generation with Pluggable TTP Framework Mapping

AgentsDGX agent

arXiv:2607.27528v1 Announce Type: cross Abstract: Threat modeling is essential for secure software development, yet manual analysis of cloud-native architectures is slow and demands scarce security ex

Tight Sample Complexity for Low-Rank Adaptation: Matching Bounds and Rank Selection

Model ReleasesDGX agent

arXiv:2607.27680v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) has become the standard mechanism for fine-tuning large pretrained models, yet its statistical properties remain only parti

(Towards) Scalable Reliable Automated Evaluation with Large Language Models

ResearchDGX agent

arXiv:2607.28282v1 Announce Type: new Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive. Existing automated

Towards Structurally Explainable Machine-Generated Text Detection: A Graph-Perspective Framework

ResearchDGX agent

arXiv:2505.12507v2 Announce Type: replace Abstract: Despite the success of machine-generated text detectors, the black-box nature remains a critical limitation. Traditional explainability methods rely

Training Skills Like Parameters via Self-Supervised Semantic Diffusion

AgentsDGX agent

arXiv:2607.27557v1 Announce Type: new Abstract: While Large Language Models (LLMs) demonstrate remarkable general instruction-following capabilities, they often fall short of human experts in highly s

TriShield: Zero-Utility-Loss Defense Against Privacy Backdoors in Federated Language Model Fine-Tuning via Orthogonal Gradient Projection and Optimizer State Entanglement

Model ReleasesDGX agent

arXiv:2607.27940v1 Announce Type: cross Abstract: Federated fine-tuning of large language models (LLMs) enables collaborative training without exposing raw data. However, a recent attack, NeuroImprint

Understanding Is Done Early: A Depth Division of Labor in Large Language Models and Its Use for Unbounded-Context Memory

HardwareDGX agent

arXiv:2607.28263v1 Announce Type: new Abstract: Transformer depth is not used uniformly: lower and middle layers build semantic representations, while upper layers increasingly specialize them for pre

Using Large Language Models for Idea Generation in Innovation

Model ReleasesDGX agent

arXiv:2607.27553v1 Announce Type: cross Abstract: This research evaluates the efficacy of large language models (LLMs) in generating new product ideas. To do so, we compare three pools of ideas for ne

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation

SafetyDGX agent

arXiv:2607.28590v1 Announce Type: cross Abstract: Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view t

Where and When to Commit: Candidate-Aware Decoding for Diffusion Language Models

ResearchDGX agent

arXiv:2607.28166v1 Announce Type: new Abstract: Diffusion language models (DLMs) expose a provisional prediction at every denoising step, creating an opportunity for generation-time early exit that st

WIDE: Boosting Adaptive LLM Inference via Token-level Dynamic Width Pruning

ApplicationsDGX agent

arXiv:2607.28418v1 Announce Type: cross Abstract: Pruning is a promising approach for improving the efficiency of LLMs. Existing static structured pruning methods are hardware-friendly and can deliver

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

Model ReleasesDGX agent

arXiv:2607.28478v1 Announce Type: new Abstract: As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in

30 Jul 2026

A large-scale corpus of religious radio broadcast transcripts from webstream recordings in the United States

ResearchDGX agent

arXiv:2607.26249v1 Announce Type: new Abstract: Religious radio is a widespread but understudied form of mass communication in the United States, and content-level analysis of it has been constrained

AdaMARP: An Adaptive Multi-Agent Interaction Framework for General Immersive Role-Playing

Model ReleasesDGX agent

arXiv:2601.11007v2 Announce Type: replace-cross Abstract: LLM role-playing aims to portray arbitrary characters in interactive narratives, yet existing systems often suffer from limited immersion and

AgentGUI: An Interface for Observing and Steering Long-Running AI Agents

Local AiDGX agent

arXiv:2607.26300v1 Announce Type: new Abstract: AI agents are increasingly adept at tackling complex, long-running tasks. With the rapid surge of autonomous capabilities, human oversight is systematic

AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents

SafetyDGX agent

arXiv:2607.26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by

Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach

Model ReleasesDGX agent

arXiv:2607.26317v1 Announce Type: cross Abstract: Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a

APEX-Accounting

Model ReleasesDGX agent

arXiv:2607.27189v1 Announce Type: new Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountant

ARC-Encoder: learning compressed text representations for large language models

Model ReleasesDGX agent

arXiv:2510.20535v2 Announce Type: replace Abstract: Recent techniques such as retrieval-augmented generation or chain-of-thought reasoning have led to longer contexts and increased inference costs. Co

ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM

ResearchDGX agent

arXiv:2506.14766v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) frequently hallucinate by over-committing to spurious visual cues. Prior remedies-Visual and Instruct

AtmosERC: Modeling Dialogue-Level Affective Atmosphere for Emotion Recognition in Conversation

ResearchDGX agent

arXiv:2607.26726v1 Announce Type: new Abstract: Emotion Recognition in Conversation (ERC) aims to predict utterance-level emotions in dialogues and has largely advanced through context-centric modelin

Automated Multilabel Mpox Research Classification with Explainable Transformer Models

ApplicationsDGX agent

arXiv:2607.26700v1 Announce Type: new Abstract: The Mpox outbreak remains a serious public health issue, with the WHO (World Health Organization) reporting increasing cases in some regions. Research o

Characterizing Human-Likeness in AI Generated Poetry: A Zero-shot Classification Study

ResearchDGX agent

arXiv:2607.26221v1 Announce Type: new Abstract: With the advancement of AI technologies, Generative AI (GenAI) and human written text have become nearly indistinguishable. Additionally, the global sta

ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information

ResearchDGX agent

arXiv:2106.16038v3 Announce Type: replace Abstract: Recent pretraining models in Chinese neglect two important aspects specific to the Chinese language: glyph and pinyin, which carry significant synta

Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting

Model ReleasesDGX agent

arXiv:2607.26200v1 Announce Type: new Abstract: Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag. We evalu

CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG

Model ReleasesDGX agent

arXiv:2607.26470v1 Announce Type: new Abstract: Multi-turn information-seeking conversations require both multi-hop reasoning and long-range dependency tracking across turns. However, existing RAG sys

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

ResearchDGX agent

arXiv:2607.26179v1 Announce Type: cross Abstract: LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human c

Constitutional Midtraining: Content Presence Drives Alignment Gains

SafetyDGX agent

arXiv:2607.26654v1 Announce Type: new Abstract: Post-training alignment is often shallow, eroding under fine-tuning. Whether midtraining interventions, cleanly isolated from post-training, can produce

Contrastive ESA: Human Evaluation of Multiple Translations at Once

ResearchDGX agent

arXiv:2607.26640v1 Announce Type: new Abstract: Current human evaluation of machine translation typically assesses single outputs in isolation, a paradigm that suffers from high annotator noise and co

Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?

Model ReleasesDGX agent

arXiv:2607.26952v1 Announce Type: new Abstract: We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

Model ReleasesDGX agent

arXiv:2607.27178v1 Announce Type: new Abstract: State-of-the-art retrieval models increasingly rely on closed training data, creating a reproducibility gap. We present an open end-to-end recipe for tr

Diagnosing Fine-Grained Inconsistency Classification in Financial Disclosure Text

Model ReleasesDGX agent

arXiv:2607.26368v1 Announce Type: new Abstract: Financial disclosures contain numerical claims, temporal statements, entity references, policy commitments, and risk descriptions that may conflict in q

DialectLLM: A Dialect-Aware Dialog[ue] Generation Framework Beyond Standard American English

Model ReleasesDGX agent

arXiv:2601.22888v4 Announce Type: replace Abstract: More than 80% of the 1.6B English speakers do not use Standard American English (SAE), yet LLMs often fail to correctly identify non-SAE dialects an

Dice Loss for Data-imbalanced NLP Tasks

ResearchDGX agent

arXiv:1911.02855v5 Announce Type: replace Abstract: Many NLP tasks such as tagging and machine reading comprehension are faced with the severe data imbalance issue: negative examples significantly out

DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models

SafetyDGX agent

arXiv:2607.26891v1 Announce Type: new Abstract: Sequence labeling is a fine-grained information extraction task, yet existing large language model-based approaches suffer from insufficient domain alig

Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens

ResearchDGX agent

arXiv:2607.26350v1 Announce Type: cross Abstract: Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be use

Do Methods Support the Claims? Intra-Paper Verification for Peer Review

SafetyDGX agent

arXiv:2607.26066v1 Announce Type: new Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review. Existing automated nove

← Previous
1…1314151617…128
Next →