AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
61+ results
16 Apr 2026

Synthesis: Arxiv-Cs-Cl

SynthesesDGX agent

Auto-generated synthesis of 505 entries about arxiv-cs-cl

12 Aug 2026

Actions Speak Louder than Words: Measuring Cross-Lingual Policy Retention in Tool-Using Agents

SafetyDGX agent

arXiv:2608.11110v1 Announce Type: new Abstract: When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compa

ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS

ResearchDGX agent

arXiv:2608.10606v1 Announce Type: new Abstract: ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading err


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

Assessing Reliability of BERT-Based Models on Question Answering Tasks

ApplicationsDGX agent

arXiv:2608.10806v1 Announce Type: new Abstract: Reliability estimation of large language models is in many cases as crucial as their accuracy, as reliable models are more trustworthy, robust, and suit

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

SafetyDGX agent

arXiv:2512.06227v3 Announce Type: replace Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis

Bayesian-Agent: Posterior-Guided Skill Evolution Across LLM Agent Harnesses

Model ReleasesDGX agent

arXiv:2606.08348v2 Announce Type: replace Abstract: LLM agents increasingly rely on prompts, tools, memory, SOPs, skills, and harness feedback, yet current self-evolution pipelines often update these

Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders

ResearchDGX agent

arXiv:2608.11197v1 Announce Type: cross Abstract: Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structur

Calibrating Post-Training Feature Shifts for LLM Data Contamination Detection

ResearchDGX agent

arXiv:2608.10462v1 Announce Type: new Abstract: Large language models (LLMs) are trained on massive and largely undisclosed corpora that may contain copyrighted or privacy-sensitive content. Data cont

Can Released LLM Vocabularies Support Token-Level Estimation of Hidden Corpora?

Model ReleasesDGX agent

arXiv:2608.10690v1 Announce Type: new Abstract: Pretraining corpus composition shapes LLM capabilities, but it often remains hidden even when model weights are released. Prior work has inferred corpus

Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory

SafetyDGX agent

arXiv:2608.09937v1 Announce Type: new Abstract: Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers d

Certify or Refuse: A Cross-Model Map for Selective Risk Control with Coverage Floors under Covariate Shift

Local AiDGX agent

arXiv:2608.10893v1 Announce Type: new Abstract: Certified selective predictors attain whatever coverage they attain; operators impose an automation floor: answer at least a eta-fraction of shifted tar

Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design

AgentsDGX agent

arXiv:2608.10299v1 Announce Type: new Abstract: Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, s

Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025

ResearchDGX agent

arXiv:2608.09936v1 Announce Type: new Abstract: Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We e

ConRub-Med: Reinforcement Learning with Consensus Rubrics for Open-Ended Medical Question Answering

SafetyDGX agent

arXiv:2608.10996v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards has been especially effective in mathematics and coding, where answers can be checked automatically. Many

Cost-Efficient Estimation of General Abilities Across Benchmarks

Model ReleasesDGX agent

arXiv:2604.01418v2 Announce Type: replace Abstract: Thousands of diverse benchmarks have been developed to measure the quality of large language models (LLMs). Yet prior work has demonstrated that LLM

Cracks in the Foundation: Seemingly Minor Architectural Choices Impact Long Context Extension

Model ReleasesDGX agent

arXiv:2608.10296v1 Announce Type: new Abstract: One might imagine that architectural variations within the dense transformer paradigm have a limited effect on accuracy. However, we demonstrate that th

Data Attribution of Emergent Misalignment with Persona Features

SafetyDGX agent

arXiv:2608.11025v1 Announce Type: new Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leadi

Decomposition-Induced Context-Memory Conflict: When Fact-Checking Pipelines Contradict Their Own Source Text

ResearchDGX agent

arXiv:2608.10627v1 Announce Type: new Abstract: Decompose-then-verify pipelines, including FActScore-style fact-checkers and long-form factuality evaluators, first split a passage into atomic claims b

Detecting an Effect Is Not Learning to Act on It: A Reward-SNR Floor for LLM Acquisition Agents

SafetyDGX agent

arXiv:2608.10441v1 Announce Type: cross Abstract: Many pipelines can pay a per-example cost to acquire an auxiliary, model-derived observation -- an LLM's structured reasoning, a slow oracle, an expen

DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student Distillation

Model ReleasesDGX agent

arXiv:2608.10636v1 Announce Type: cross Abstract: Visual document retrieval (VDR) is dominated by multi-billion-parameter models that are slow to index at full corpus scale and expensive to serve. Pri

Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

Model ReleasesDGX agent

arXiv:2608.10008v1 Announce Type: cross Abstract: LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate;

Dual-Loop Self-Evolution via Verifiable Emotion Feedback for Multi-Turn Empathetic Dialogue

SafetyDGX agent

arXiv:2608.10626v1 Announce Type: new Abstract: Large language models have demonstrated conversational capabilities, yet empathetic competence remains challenging. Empathetic support is inherently mul

Enhancing Automated Essay Scoring With Three Techniques: Two-Stage Fine-Tuning, Score Alignment, and Self-Training

SafetyDGX agent

arXiv:2602.01747v2 Announce Type: replace Abstract: Automated Essay Scoring (AES) plays a crucial role in education by providing scalable and efficient assessment tools. However, in real-world setting

Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases

SafetyDGX agent

arXiv:2608.10503v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NL

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

ResearchDGX agent

arXiv:2608.10698v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chine

Faster Superword Tokenization

ResearchDGX agent

arXiv:2604.05192v2 Announce Type: replace Abstract: Byte Pair Encoding (BPE) is a widely used tokenization algorithm, whose tokens cannot extend across pre-tokenization boundaries, functionally limiti

FoR-SALE: Frame of Reference-guided Spatial Adjustment in LLM-based Diffusion Editing

SafetyDGX agent

arXiv:2509.23452v2 Announce Type: replace-cross Abstract: Current text-to-image generation models, even state-of-the-art models, exhibit a significant performance gap when spatial expressions are desc

Generation-Step-Aware Framework for Cross-Modal Representation and Control in Multilingual Speech-Text Models

SafetyDGX agent

arXiv:2601.17387v3 Announce Type: replace Abstract: Multilingual speech-text models rely on cross-modal language alignment to transfer knowledge between speech and text, but it remains unclear whether

How Robust Are LLMs to Vietnamese Dialects?

Model ReleasesDGX agent

arXiv:2608.10414v1 Announce Type: new Abstract: Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects th

InSight-doc: Agentic Visual Perception for Long-Document Understanding

Model ReleasesDGX agent

arXiv:2608.10628v1 Announce Type: cross Abstract: Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we

Leveraging Human Reading Behavior for Keyphrase Extraction: A Webcam-based Eye-tracking Corpus

ResearchDGX agent

arXiv:2608.10688v1 Announce Type: new Abstract: Purpose: Keyphrases are statistically and semantically important textual units that can also attract readers' attention during comprehension. However, e

Mapping and Measuring the Behavioral Evolution of Large Language Models

Model ReleasesDGX agent

arXiv:2608.11027v1 Announce Type: cross Abstract: Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across genera

Mitigating Context Interference for Reliable and Efficient Search Agents

AgentsDGX agent

arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are s

MT-PingEval: Evaluating Multi-Turn Collaboration with Private Information Games

AgentsDGX agent

arXiv:2602.24188v2 Announce Type: replace Abstract: We present a scalable and verifiable methodology for evaluating language models in multi-turn interactions, using a suite of collaborative games tha

Multilingual Embedding Probes Fail to Generalize Across Learner Corpora

ResearchDGX agent

arXiv:2604.07095v2 Announce Type: replace Abstract: Do multilingual embedding models encode a language-general representation of proficiency? We investigate this by training linear and non-linear prob

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

Local AiDGX agent

arXiv:2608.11167v1 Announce Type: cross Abstract: Existing Multimodal Large Language Models (MLLMs) predominantly rely on image-text pairs for modality alignment pretraining, mapping global image repr

MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales

ResearchDGX agent

arXiv:2608.10974v1 Announce Type: new Abstract: Scientific papers contain fine-grained records of problem solving: authors mention technical obstacles and methods that were used to address them, often

myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

Model ReleasesDGX agent

arXiv:2608.11036v1 Announce Type: new Abstract: Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work prese

No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding

Model ReleasesDGX agent

arXiv:2503.05061v3 Announce Type: replace Abstract: Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as bus

Off-Axis, On Purpose: Where a Transformer Computes Concepts and Why it Does So

ResearchDGX agent

arXiv:2608.10251v1 Announce Type: new Abstract: A transformer's answer lives on one axis: the direction its unembedding reads. Its intermediate states largely do not, and that off-axis position is usu

OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents

AgentsDGX agent

arXiv:2608.09988v1 Announce Type: cross Abstract: Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can b

Order Matters: LVLMs as Judges for Temporal Reasoning in Image Sequences

ResearchDGX agent

arXiv:2608.10908v1 Announce Type: cross Abstract: As generative multimedia evolves from static image synthesis to complex, interleaved visual narratives, a foundational bottleneck has emerged: the jud

PERCEPT: A Corpus for POS Tagging and Analysis of Persian-English Code-Mixing

ResearchDGX agent

arXiv:2608.10109v1 Announce Type: new Abstract: Social media has become a major venue for multilingual communication, where users frequently mix multiple languages within a single utterance. Although

Position Encoding in Transformers: From Absolute and Relative Methods to Rotary Position Embeddings and Long-Context Scaling

SafetyDGX agent

arXiv:2608.10021v1 Announce Type: new Abstract: Self-attention models content-dependent interactions between tokens but does not by itself encode token order. Position encoding addresses this limitati

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference

Model ReleasesDGX agent

arXiv:2608.10288v1 Announce Type: cross Abstract: The Large Language Model from Power Law Decoder Representations (PLDR-LLM) and its attention, Power Law Graph Attention (PLGA), replace the fixed bili

REAP: Relation-Aware Elicitation and Parsing for Closed-Book Knowledge Base Construction from LLMs

Model ReleasesDGX agent

arXiv:2608.10963v1 Announce Type: new Abstract: We present the REAP system for the AKBC Shared Task 2026 on constructing knowledge bases from language models in a closed-book setting, subject to a bud

Reinforcement Learning-based Semi-supervised Knowledge Distillation with LLM-as-a-Judge

ResearchDGX agent

arXiv:2604.02621v2 Announce Type: replace Abstract: Reinforcement Learning (RL) substantially improves the reasoning capabilities of language models, but most existing RL fine-tuning approaches rely e

ReRound: Reconstructive Rounding to Resolve Midpoint Ambiguity in Calibration-Free LLM Quantization

Model ReleasesDGX agent

arXiv:2608.11045v1 Announce Type: cross Abstract: ReRound (Reconstructive Rounding) is a post-training quantization method that addresses the midpoint ambiguity inherent in standard round-to-nearest (

Robust and Secure Code Watermarking for Large Language Models via ML/Crypto Codesign

ResearchDGX agent

arXiv:2502.02068v3 Announce Type: replace-cross Abstract: This paper introduces RoSeMary, the first-of-its-kind ML/Crypto codesign watermarking framework that regulates LLM-generated code to avoid int

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

Model ReleasesDGX agent

arXiv:2608.10670v1 Announce Type: new Abstract: At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-reso

Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching

ApplicationsDGX agent

arXiv:2608.11030v1 Announce Type: cross Abstract: Patent retrieval and matching based on large language models (LLMs) play a vital role in intellectual property protection. However, due to the complex

Share First, Route What Remains: A Unified Framework for Token-Adaptive MoE Computation

ResearchDGX agent

arXiv:2608.10392v1 Announce Type: cross Abstract: Mixture-of-experts (MoE) models have recently moved beyond routing a fixed number of complete experts. Shared-expert designs preserve reusable knowled

Simplex Relaxation for Discrete Diffusion

ResearchDGX agent

arXiv:2608.10615v1 Announce Type: new Abstract: Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associate

StreamFlow: Dynamic Memory Flows for Streaming Video Understanding

ResearchDGX agent

arXiv:2608.10949v1 Announce Type: cross Abstract: Streaming video understanding requires multimodal large language models (MLLMs) to preserve relevant evidence from continuously evolving streams under

TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification

ResearchDGX agent

arXiv:2608.11044v1 Announce Type: new Abstract: Hierarchical Text Classification (HTC), as a critical text mining task, faces challenges such as complex label hierarchies and class imbalance. Existing

TemMed-Bench: Evaluating Temporal Medical Image Reasoning in Vision-Language Models

Model ReleasesDGX agent

arXiv:2509.25143v2 Announce Type: replace-cross Abstract: Existing medical reasoning benchmarks for vision-language models primarily focus on analyzing a patient's condition based on an image from a s

Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance

SafetyDGX agent

arXiv:2608.11008v1 Announce Type: new Abstract: Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans,

The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues

SafetyDGX agent

arXiv:2603.20907v5 Announce Type: replace Abstract: As users increasingly turn to LLMs for practical and personal advice, they become vulnerable to subtle steering toward hidden incentives misaligned

The Illusion of Cross-Lingual Safety in Low-Resource Languages

Local AiDGX agent

arXiv:2608.11146v1 Announce Type: new Abstract: Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. How

The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs

Model ReleasesDGX agent

arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degrada

← Previous
1
Next →
7,646 results
← Previous
123…128
Next →