AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
10 Jun 2026

ArabiGEE: A Hierarchical Taxonomy for Arabic Grammatical Error Explanation

ResearchDGX agent

arXiv:2606.10765v1 Announce Type: new Abstract: We introduce ArabiGEE, the first comprehensive Arabic grammatical error explanation (GEE) taxonomy grounded in explicit error types. Unlike existing GEE

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

ResearchDGX agent

arXiv:2606.10657v1 Announce Type: new Abstract: Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes t

Attention Amnesia in Hybrid LLMs: When CoT Fine-Tuning Breaks Long-Range Recall, and How to Fix It

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2606.11052v1 Announce Type: new Abstract: Chain-of-thought (CoT) supervised fine-tuning (SFT) is widely adopted to improve reasoning ability, yet we find that it systematically degrades long-con

Automated Alignment between Elicitation Interviews and Requirements

SafetyDGX agent

arXiv:2510.08622v2 Announce Type: replace Abstract: Software requirements are derived from a variety of elicitation techniques, many of which have a conversational nature, like interviews. However, ev

Automated Scoring of Arabic Text Using Large Language Models: A Literature Review

SafetyDGX agent

arXiv:2606.09830v1 Announce Type: new Abstract: In modern educational systems, Automatic Text Scoring (ATS) plays a central role by enabling scalable and consistent evaluation of learner responses wit

Benchmarking and Exploring the Capabilities of LLMs for Attack Investigations

Model ReleasesDGX agent

arXiv:2606.10281v1 Announce Type: cross Abstract: This paper presents AuditBench, a new benchmark dataset for evaluating the capabilities of LLMs at investigating security-related system audit logs. W

BenSyc: Benchmarking Conversational Sycophancy and Human Alignment in LLMs for Bengali Contexts

Model ReleasesDGX agent

arXiv:2606.10061v1 Announce Type: new Abstract: Large language models (LLMs) increasingly participate in emotionally sensitive social conversations, where responses may shift from balanced support tow

Beyond Memorization: Distinguishing Between Pattern-Based and Epistemic Reasoning in LLMs Using Epistemic Puzzles

Model ReleasesDGX agent

arXiv:2603.21350v2 Announce Type: replace Abstract: Epistemic reasoning requires agents to infer the state of the world from partial observations and information about other agents' knowledge. Prior w

CGES: Confidence-Guided Early Stopping for Efficient and Accurate Self-Consistency

ResearchDGX agent

arXiv:2511.02603v2 Announce Type: replace Abstract: Large language models (LLMs) are often queried multiple times at test time, with predictions aggregated by majority vote. While effective, this self

CodeAlchemy: Synthetic Code Rewriting at Scale

Model ReleasesDGX agent

arXiv:2606.10087v1 Announce Type: new Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats. While synthetic data has proven transformative f

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Model ReleasesDGX agent

arXiv:2601.18026v2 Announce Type: replace Abstract: Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, e

Compiling Rewrite Rules to Finite-State Transducers with the Worsening Trick

ApplicationsDGX agent

arXiv:2606.10059v1 Announce Type: cross Abstract: Finite-state transducers (FSTs) are essential for modeling string rewriting in computational linguistics and natural language processing (NLP), partic

Continual LLM Upcycling: A Predictor-Gated Bank-Wise Sparsity Training Recipe for Dense-to-Sparse LLMs

Model ReleasesDGX agent

arXiv:2606.10722v1 Announce Type: new Abstract: We study dense-to-sparse continual training as a way to construct channel-sparse large language models from dense checkpoints. Starting from a Qwen2.5-8

ConvMemory v2: A Recall-Preserving Top-10 Evidence Reranker for Conversational Memory Retrieval

Model ReleasesDGX agent

arXiv:2606.10842v1 Announce Type: new Abstract: We describe ConvMemory v2, an opt-in token-evidence reranker that sits after the lightweight ConvMemory v1 reranker and reorders only v1's protected top

CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback

Model ReleasesDGX agent

arXiv:2504.02323v4 Announce Type: replace Abstract: Large language models (LLMs) have created new opportunities to assist teachers and support student learning. While researchers have explored various

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories

AgentsDGX agent

arXiv:2606.11176v1 Announce Type: cross Abstract: Data tells stories that shape society; the data journalist's job is to turn raw information into stories non-experts can trust. A high-quality news fe

DECSELFMASK: Leveraging Unlabeled Text via Self-Relevance-Guided Masking for Decoder-Only Classification

ResearchDGX agent

arXiv:2606.09466v2 Announce Type: replace Abstract: Classification tasks require annotated data, which can often be expensive, time-consuming, or even unfeasible to collect. This is the case of the me

Density Field State Space Models: 1-Bit Distillation, Efficient Inference, and Knowledge Organization in Mamba-2

Local AiDGX agent

arXiv:2606.10932v1 Announce Type: new Abstract: We present Density Field State Space Models (DF-SSM), a framework for compressing SSMs to a 1-bit scaffold with int8 low-rank correction. Applied to Mam

Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark

Model ReleasesDGX agent

arXiv:2606.10400v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors,

Does Reasoning Preserve Alignment? On the Trustworthiness of Large Reasoning Models

SafetyDGX agent

arXiv:2606.11046v1 Announce Type: new Abstract: Instruction-tuned LLMs are increasingly converted into reasoning models through post-training to improve multi-step task performance. This conversion is

Early-Token Confidence Predicts Reasoning Quality in Multi-Agent LLM Debate

SafetyDGX agent

arXiv:2606.10307v1 Announce Type: new Abstract: Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whethe

Enhancing Multilingual LLM-based ASR with Mixture of Experts and Dynamic Downsampling

SafetyDGX agent

arXiv:2606.10439v1 Announce Type: cross Abstract: The rapid progress of large language models (LLMs) has opened up a new frontier for automatic speech recognition (ASR), making their effective integra

Entropy, Disagreement, and the Limits of Foundation Models in Genomics

TutorialsDGX agent

arXiv:2604.04287v2 Announce Type: replace-cross Abstract: Foundation models in genomics have shown mixed success compared to their counterparts in natural language processing. Yet, the reasons for the

From Genes to Tokens: a GWAS-inspired Approach for Interpretable Stylometric Analysis

ResearchDGX agent

arXiv:2606.09543v2 Announce Type: replace Abstract: This short paper introduces a stylometric interpretation method inspired by genome-wide association studies (GWAS). Each 'gene' token's association

From Observation to Intervention: A Causal Audit of Expert Importance in Mixture-of-Experts Models

Model ReleasesDGX agent

arXiv:2606.10703v1 Announce Type: cross Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of target

Generative Archetype-Grounded Item Representations for Sequential Recommendation

ApplicationsDGX agent

arXiv:2606.11023v1 Announce Type: cross Abstract: Sequential recommendation aims to predict users' next interaction with items by analyzing their historical behavior. However, the limited quality of i

GhazalBench: Evaluating LLM Understanding and Canonical Surface-Form Access in Persian Ghazals

Model ReleasesDGX agent

arXiv:2603.09979v2 Announce Type: replace Abstract: Persian poetry plays an active role in Iranian cultural practice, where verses by canonical poets such as Hafez are frequently quoted, paraphrased,

How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs

ResearchDGX agent

arXiv:2606.10646v1 Announce Type: cross Abstract: Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat

inversedMixup: Data Augmentation via Inverting Mixed Embeddings

ResearchDGX agent

arXiv:2601.21543v3 Announce Type: replace Abstract: Mixup generates augmented samples by linearly interpolating inputs and labels with a controllable ratio. However, since it operates at the latent em

It Takes One to Bias Them All: Breaking Bad with One-Shot GRPO

SafetyDGX agent

arXiv:2606.10931v1 Announce Type: new Abstract: Warning: This paper contains several toxic and offensive statements. Modern large language models (LLMs) are typically aligned through large-scale post-

KCSAT-ML: Probing Reasoning Models with Nationwide-Cohort Human Difficulty

Model ReleasesDGX agent

arXiv:2606.10403v1 Announce Type: new Abstract: Math reasoning benchmarks have proliferated, yet most lack a per-item difficulty signal grounded in actual human performance. We introduce KCSAT-ML, a d

Large Language Models as Modal Models in Linguistics

ResearchDGX agent

arXiv:2606.10467v1 Announce Type: new Abstract: The rapid advancement of large language models (LLMs) has intensified debates about their significance for linguistic theory. These debates are commonly

Leveraging Social Media Data for COVID-19 Studies

ResearchDGX agent

arXiv:2606.10459v1 Announce Type: cross Abstract: Nowadays, social media networks have become widely preferred sources of information. Especially during the time of the Coronavirus disease 2019 COVID

Lightweight Latent Reasoning for Narrative Tasks

SafetyDGX agent

arXiv:2512.02240v2 Announce Type: replace Abstract: Large language models (LLMs) tackle complex tasks by generating long chains of thought or 'reasoning traces' that act as latent variables in the gen

Measuring Human Value Expression in Social Media Texts: Calibrated LLM Annotation and Encoder Transfer

SafetyDGX agent

arXiv:2606.11018v1 Announce Type: new Abstract: Measuring subjective constructs in naturally occurring social media text requires annotation procedures that are theoretically grounded, empirically val

Mechanistic Analysis of Alignment Algorithms in Language Models

SafetyDGX agent

arXiv:2606.09850v1 Announce Type: cross Abstract: Post-training alignment algorithms are predominantly evaluated as black boxes, obscuring how they reshape language models' internal computations. We p

MIRAGE: A Polarity-Flipping Encoding Subspace in LLM Agents

Model ReleasesDGX agent

arXiv:2606.10304v1 Announce Type: new Abstract: When LLM agents are coerced into covertly encoding sensitive data (Base64, ROT13, acrostic, synonym chains, and beyond), the resulting outputs evade out

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

SafetyDGX agent

arXiv:2606.11167v1 Announce Type: new Abstract: Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current

Multilingual Word-Level Forced Alignment with Self-Supervised Representations and Learned Dynamic Programming

SafetyDGX agent

arXiv:2606.10675v1 Announce Type: new Abstract: We present a method for accurate multilingual word-level forced alignment, consisting of an alignment encoder and a learned alignment decoder. The encod

N-GRPO: Embedding-Level Neighbor Mixing for Enhanced Policy Optimization

Model ReleasesDGX agent

arXiv:2606.10768v1 Announce Type: cross Abstract: The success of Large Language Models in mathematical reasoning relies heavily on the generation of diverse and valid solution paths during the rollout

Open Korean Corpora: A Practical Report

ResearchDGX agent

arXiv:2012.15621v3 Announce Type: replace Abstract: Korean is often referred to as a low-resource language in the research community. While this claim is partially true, it is also because the availab

OpenRTLSet: A Fully Open-Source Dataset for Large Language Model-based Verilog Module Design

Model ReleasesDGX agent

arXiv:2606.10285v1 Announce Type: new Abstract: OpenRTLSet introduces the largest fully open-source dataset for hardware design, offering over 131,000 diverse Verilog code samples to the research comm

PADD: Path-Aligned Decompression Distillation for Non-Router Teacher to Guide MoE Student Learning

SafetyDGX agent

arXiv:2606.10369v1 Announce Type: new Abstract: As large language models (LLMs) continue to scale, it becomes increasingly challenging to grow model capacity under fixed computation budgets. We propos

ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language Models

SafetyDGX agent

arXiv:2606.10581v1 Announce Type: new Abstract: Speech carries more information than just words: a child's voice, a fearful tone, or a noisy background should all lead a sufficiently competent spoken-

Parallel Causal Associative Fields: Gated Sparse Memory for Long-Context Language Modeling

Model ReleasesDGX agent

arXiv:2606.10435v1 Announce Type: cross Abstract: Transformers achieve strong language modeling performance by providing direct token-to-token communication paths, but causal self-attention scales qua

Parametric Knowledge is Not All You Need: Toward Honest Large Language Models via Retrieval of Pretraining Data

Model ReleasesDGX agent

arXiv:2601.21218v2 Announce Type: replace Abstract: Large language models (LLMs) are highly capable of answering questions, but they are often unaware of their own knowledge boundary, i.e., knowing wh

Pre-AF 13: An Interpretable Atrial Fibrillation Risk Score Mined from Discharge Reports

ResearchDGX agent

arXiv:2606.10725v1 Announce Type: cross Abstract: Background. Atrial fibrillation (AF) is the most prevalent cardiac arrhythmia and a major determinant of prognosis. Established AF risk scores rely on

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models

ResearchDGX agent

arXiv:2606.10537v1 Announce Type: new Abstract: Diffusion large language models (dLLMs) re-encode the entire prefix at every denoising step, causing recomputation that scales quadratically with contex

ProbeLLM: Automating Principled Diagnosis of LLM Failures

Model ReleasesDGX agent

arXiv:2602.12966v2 Announce Type: replace Abstract: Understanding how and why large language models (LLMs) fail is becoming a central challenge as models rapidly evolve and static evaluations fall beh

Pushing the Limits of LLM Tool Calling via Experiential Knowledge Integration and Activation

AgentsDGX agent

arXiv:2606.10875v1 Announce Type: new Abstract: Large language models (LLMs) rely on tool use to act as autonomous agents, yet often fail in multi-step execution due to insufficient tool-related knowl

REAL: A Reasoning-Enhanced Graph Framework for Long-Term Memory Management of LLMs

Model ReleasesDGX agent

arXiv:2606.10694v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly expected to interact with users over long time horizons. However, due to their finite context window, LLMs

Recovering the Zipfian Distribution in Unsupervised Term Discovery

SafetyDGX agent

arXiv:2606.10781v1 Announce Type: cross Abstract: Unsupervised term discovery involves segmenting unlabelled speech into word- or syllable-like units and clustering these into a lexicon of candidate t

RedAct: Redacting Agent Capability Traces for Procedural Skill Protection

Model ReleasesDGX agent

arXiv:2606.10813v1 Announce Type: cross Abstract: Users rely on execution traces to observe agent behavior, diagnose failures, and ensure accountability. These traces contain rich procedural detail, i

Representation-Aware Advantage Estimation: Your Reward Model Provides More Than A Scalar Output

ResearchDGX agent

arXiv:2606.10528v1 Announce Type: cross Abstract: Current reinforcement learning from human feedback (RLHF) methods primarily rely on scalar rewards from a trained reward model (RM). While effective,

Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages

Model ReleasesDGX agent

arXiv:2510.07061v2 Announce Type: replace Abstract: While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated a

Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster

ResearchDGX agent

arXiv:2603.07238v2 Announce Type: replace Abstract: Similarities between language representations derived from Self-Supervised Speech Models (S3Ms) have been observed to primarily reflect geographic p

Selection, Not Salience: The Shape and Limits of Personalization in Social Highlighting

SafetyDGX agent

arXiv:2606.10398v1 Announce Type: cross Abstract: Does personalizing what a reader sees pay off, and where does it stop? Using a social web highlighter and a co-readership identity control (the same d

Small Data, Big Noise: Adversarial Training for Robust Parameter-Efficient Fine-Tuning

Model ReleasesDGX agent

arXiv:2606.10610v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become essential for adapting foundation models to downstream NLP tasks. However, current PEFT methods often

Speaker Group Encoding in Self-supervised Speech Recognition Models

SafetyDGX agent

arXiv:2606.10654v1 Announce Type: new Abstract: We investigate what self-supervised speech recognition models (S3Ms) learn about speaker groups (SGs). We examine several states of S3Ms: pretrained, fi

SpeechJBB: Probing Safety Alignment and Comprehension in Large Audio Language Models under Code-Switched Speech

SafetyDGX agent

arXiv:2606.06037v2 Announce Type: cross Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evaluated on m

← Previous
1…3839404142…129
Next →