AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
30 Jun 2026

LatentRevise: Learning from Zero-Hit Reasoning

SafetyDGX agent

arXiv:2606.29938v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by hard prompts on which correct trajectories have low probability, so sampling mi

Legal Domain Adaptation of Modern BERT Models

ApplicationsDGX agent

arXiv:2606.28538v1 Announce Type: new Abstract: We investigate domain adaptation of modern BERT models in the legal domain. We further pre-train ModernBERT on all US court opinions using the masked la

LLM Agents Are Latent Context Managers: Eliciting Self-Managed Context via a Proprioceptive Dashboard

Model ReleasesDGX agent

arXiv:2606.30005v1 Announce Type: new Abstract: Long-horizon tool agents are bottlenecked by how their context grows toward the limits of the context window. Recent systems make context management age


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue

ResearchDGX agent

arXiv:2509.02292v2 Announce Type: replace Abstract: What if large language models could not only infer human mindsets but also expose every blind spot in team dialogue such as discrepancies in the tea

MaDI-Bench: An End-to-End Data Integration Benchmark

Model ReleasesDGX agent

arXiv:2606.30371v1 Announce Type: cross Abstract: Data integration combines heterogeneous data sets into a single, coherent representation. Data integration involves a sequence of interdependent tasks

MAM-AI: An On-Device Medical Retrieval-Augmented Generation System for Nurses and Midwives in Zanzibar

Model ReleasesDGX agent

arXiv:2606.29580v1 Announce Type: new Abstract: Maternal and newborn mortality remain among the highest in sub-Saharan Africa, where midwifery care is often delivered by nurses who lack midwifery trai

mamabench and mamaretrieval: Benchmarks for Evaluating Medical Retrieval-Augmented Generation in Maternal, Neonatal, and Reproductive Health

Model ReleasesDGX agent

arXiv:2606.29467v1 Announce Type: new Abstract: Medical question-answering benchmarks rarely cover the maternal, neonatal, child, and reproductive-health questions a nurse-midwife asks, and, to our kn

Managing Map Cardinality in Automatic Disease Classification Mapping: Balancing Precision, Recall and Coverage

ResearchDGX agent

arXiv:2606.29750v1 Announce Type: new Abstract: Automatic mapping between disease classification systems, such as the International Classification of Diseases (ICD), is a challenging yet essential tas

Masked Diffusion Decoding as x-Prediction Flow

SafetyDGX agent

arXiv:2606.29066v1 Announce Type: new Abstract: Masked diffusion language models (MDLMs) generate text by iteratively unmasking tokens, but their standard decoder reduces each step to a binary action:

MauBERT: Universal Phonetic Inductive Biases for Few-Shot Acoustic Units Discovery

TutorialsDGX agent

arXiv:2512.19612v2 Announce Type: replace Abstract: This paper introduces MauBERT, a multilingual extension of HuBERT that leverages articulatory features for robust cross-lingual phonetic representat

MemDelta: Controlled Baselines and Hidden Confounds in Agent Memory Evaluation

Model ReleasesDGX agent

arXiv:2606.29914v1 Announce Type: new Abstract: Agent memory systems are increasingly evaluated against RAG and full-context baselines, but reported gains often mix changes in the memory method with c

Memory-Managed Long-Context Attention: A Preliminary Study of Editable Request-Local Memory

Model ReleasesDGX agent

arXiv:2606.28876v1 Announce Type: new Abstract: Long-context language models often conflate two different goals: compressing history into an efficient state, and maintaining reliable long-term memory.

MIThinker: A Plug-and-Play Policy-Optimized Thinker For Motivational Interviewing Counseling

SafetyDGX agent

arXiv:2606.29265v1 Announce Type: new Abstract: Reasoning large language models (LLMs) have recently made much progress in complex problem-solving, leveraging internal reasoning (or thought) to guide

Mitigating Batch Effects in Histopathology via Language-Mediated Robust Embedding Generation

ResearchDGX agent

arXiv:2606.28697v1 Announce Type: cross Abstract: Pathology foundation models (PFMs) have demonstrated strong potential across clinical and scientific applications, yet their performance is often hind

MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

Model ReleasesDGX agent

arXiv:2602.21608v2 Announce Type: replace Abstract: Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain s

Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders

ResearchDGX agent

arXiv:2507.23220v2 Announce Type: replace Abstract: Traditional topic models are effective at uncovering latent themes in large text collections. However, due to their reliance on bag-of-words represe

MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training

SafetyDGX agent

arXiv:2606.30406v1 Announce Type: new Abstract: Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabili

Morphing into Hybrid Attention Models

Model ReleasesDGX agent

arXiv:2606.30562v1 Announce Type: new Abstract: Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with line

Multi-Agentic System Leveraging Open-Source LLMs to Mitigate Disinformation Threats

Model ReleasesDGX agent

arXiv:2606.30259v1 Announce Type: new Abstract: In contemporary societies, the threat of disinformation has reached alarming levels, exacerbated by the proliferation of electronic communication, socia

Multi-Block Diffusion Language Models

ResearchDGX agent

arXiv:2606.29215v1 Announce Type: cross Abstract: Block Diffusion Language Models (BD-LMs) improve diffusion-based text generation with KV caching and flexible-length generation. A natural next step i

Multimodal Mathematical Reasoning with Diverse Solving Perspective

Model ReleasesDGX agent

arXiv:2507.02804v2 Announce Type: replace Abstract: Recent progress in large-scale reinforcement learning (RL) has notably enhanced the reasoning capabilities of large language models (LLMs), especial

Node-to-Neighborhood Semantic Consistency: Text-Topology Alignment for TAGs Anomaly Detection

SafetyDGX agent

arXiv:2606.30009v1 Announce Type: new Abstract: Graph anomaly detection (GAD) on text-attributed graphs (TAGs) is vital for applications such as fraud detection and academic integrity verification. Ex

Not-quite-human tastes: the stylized omnivorousness of LLM survey surrogates

Model ReleasesDGX agent

arXiv:2606.30085v1 Announce Type: new Abstract: Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produc

OLIVE: View-Augmented Latent Prediction with Waveform Reconstruction for Speech SSL

ResearchDGX agent

arXiv:2606.30356v1 Announce Type: new Abstract: We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that joi

Online Experiential Learning for Language Models

SafetyDGX agent

arXiv:2603.16856v2 Announce Type: replace Abstract: The prevailing paradigm for improving large language models relies on offline training with human annotations or simulated environments, leaving the

Open but Incompatible: A License Compatibility Analysis of Corpora for Low-Resource African Languages

ApplicationsDGX agent

arXiv:2606.28867v1 Announce Type: new Abstract: Creative Commons licenses dominate African NLP corpus releases, but their compatibility rules are rarely applied. CC-BY-SA and CC-BY-NC cannot be combin

Parametric Skills

Model ReleasesDGX agent

arXiv:2606.30015v1 Announce Type: new Abstract: Since intelligence fundamentally relies on efficient skill acquisition (Chollet, 2019), the ability to leverage skills is critical. For LLMs, skills, ma

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs

ResearchDGX agent

arXiv:2606.28898v1 Announce Type: new Abstract: Knowledge updating in pre-trained Large Language Models (LLMs) remains an important challenge. While continual training provides a potential avenue for

Phonological Perception of Sign Language Models

SafetyDGX agent

arXiv:2606.28667v1 Announce Type: new Abstract: Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement

Poller: Are LLMs Suitable for Evaluating the Poetry Understanding Task?

ResearchDGX agent

arXiv:2606.30556v1 Announce Type: new Abstract: Traditional automatic evaluation methods have been shown to be unsuitable for modern Chinese poetry because of the distinct nature of this literary genr

Preference-ASR: A Preference-Aware Test Set for Benchmarking ASR in the Era of Speech LLMs

ResearchDGX agent

arXiv:2606.29534v1 Announce Type: new Abstract: Popular ASR test sets adopt inconsistent conventions for numbers, disfluencies, entities, and casing, while standard normalizers erase the format distin

Preserving Fairness and Safety in Quantized LLMs Through Critical Weight Protection

SafetyDGX agent

arXiv:2601.12033v2 Announce Type: replace Abstract: Quantization is widely adopted to reduce the computational cost of large language models (LLMs); however, its implications for fairness and safety,

REAR: Test-time Preference Realignment through Reward Decomposition

SafetyDGX agent

arXiv:2606.30339v1 Announce Type: new Abstract: Aligning large language models (LLMs) with diverse user preferences is a critical yet challenging task. While post-training methods can adapt models to

Regime-Aware Peer Specialization for Robust RAG under Heterogeneous Knowledge Conflicts

SafetyDGX agent

arXiv:2606.30518v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) improves language models by grounding generation in external context. However, it can be fragile when the retrieved

Representational Depth of Evaluation Awareness Shifts With Scale in Open-Weight Language Models

Model ReleasesDGX agent

arXiv:2606.29196v1 Announce Type: cross Abstract: Do language models know when they are being tested? This question matters for AI safety: a model that recognises an evaluation context could alter its

Resolution Thresholds in VLM Detection of Harmful ASCII Art Across Construction Modes and Languages

SafetyDGX agent

arXiv:2606.29649v1 Announce Type: new Abstract: Large Vision-Language Models (VLMs) are increasingly deployed as content moderation tools, yet they remain vulnerable to jailbreak attacks in which harm

Revealing the Technology Development of Natural Language Processing: A Scientific Entity-Centric Perspective

ResearchDGX agent

arXiv:2606.29836v1 Announce Type: new Abstract: Most studies on technology development have been conducted from a thematic perspective, but the topics are coarse-grained and insufficient to accurately

Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent

Model ReleasesDGX agent

arXiv:2606.30616v1 Announce Type: new Abstract: We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We invest

SEAD: Competence-Aware On-Policy Distillation via Entropy-Guided Supervision

SafetyDGX agent

arXiv:2606.28562v1 Announce Type: new Abstract: On-policy distillation (OPD) has a property absent in offline distillation and RL: teacher supervision quality depends on student competence. Incoherent

See, Think, Learn: A Self-Taught Multimodal Reasoner

TutorialsDGX agent

arXiv:2512.02456v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) have achieved remarkable progress in integrating visual perception with language understanding. However, effecti

SHOVIR: A Benchmark for Evaluating Vision Shortcut Learning in Radiology Report Generation

Model ReleasesDGX agent

arXiv:2606.30201v1 Announce Type: cross Abstract: Current evaluation protocols for Vision-Language Models (VLMs) in Radiology Report Generation (RRG) rely on report-level metrics that measure lexical

Sifei at SemEval-2026 Task 8: Hybrid Retrieval and Query Rewriting for Multi-Turn RAG

ResearchDGX agent

arXiv:2606.28352v1 Announce Type: cross Abstract: Multi-turn retrieval-augmented generation (RAG) is challenging due to evolving user intent, conversational noise, and strict context limits. We propos

Smooth Scaling Laws Hide Stepwise Token Learning

Local AiDGX agent

arXiv:2606.29858v1 Announce Type: new Abstract: Language model loss follows remarkably regular scaling laws over model and data size, yet it remains unclear why the aggregate loss should exhibit a pow

SPARKLING: Balancing Signal Preservation and Symmetry Breaking for Width-Progressive Learning

ResearchDGX agent

arXiv:2602.02472v2 Announce Type: replace-cross Abstract: Progressive Learning (PL) reduces pre-training computational overhead by gradually increasing model scale. While prior work has extensively ex

Sparse Autoencoders are Capable LLM Jailbreak Mitigators

SafetyDGX agent

arXiv:2602.12418v2 Announce Type: replace-cross Abstract: Jailbreak attacks remain a persistent threat to large language model safety. We propose Context-Conditioned Delta Steering (CC-Delta), an SAE-

SpecMind: Cognitively Inspired, Interactive Multi-Turn Framework for Postcondition Inference

SafetyDGX agent

arXiv:2602.20610v3 Announce Type: replace-cross Abstract: Specifications are vital for ensuring program correctness, yet writing them manually remains challenging and time-intensive. Recent large lang

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models

Model ReleasesDGX agent

arXiv:2606.29815v1 Announce Type: new Abstract: Evaluating code large language models (Code LLMs) requires reliable detection of data leakage, where benchmark performance is artificially inflated by e

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models

SafetyDGX agent

arXiv:2510.12784v2 Announce Type: replace-cross Abstract: Recently, remarkable progress has been made in Unified Multimodal Models (UMMs), which integrate vision-language generation and understanding

Structure-Preserving Document Translation via Multi-Stage LLM Pipeline: A Case Study in Marathi

SafetyDGX agent

arXiv:2606.28796v1 Announce Type: new Abstract: Government documents in India are predominantly issued in regional languages such as Marathi, creating substantial accessibility barriers for non-native

Supporting Workflow Reproducibility by Linking Bioinformatics Tools across Papers and Executable Code

ResearchDGX agent

arXiv:2603.08195v2 Announce Type: replace Abstract: Motivation: The rapid growth of biological data has intensified the need for transparent, reproducible, and well-documented computational workflows.

SurrogateShield: Beyond Redaction for High-Utility, Privacy-Preserving LLM Interactions

Local AiDGX agent

arXiv:2606.29567v1 Announce Type: cross Abstract: LLM-based assistants transmit user queries verbatim to third-party API endpoints that lie outside the user's audit or control. When those queries cont

The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

SafetyDGX agent

arXiv:2510.06096v3 Announce Type: replace-cross Abstract: The objectives that Large Language Models (LLMs) implicitly optimize remain dangerously opaque, making trustworthy alignment and auditing a gr

The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

Model ReleasesDGX agent

arXiv:2606.28325v1 Announce Type: cross Abstract: Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a se

The Effect of Scripts and Formats on LLM Numeracy

ResearchDGX agent

arXiv:2601.15251v2 Announce Type: replace Abstract: Large language models (LLMs) have achieved impressive proficiency in basic arithmetic, rivaling human-level performance on standard numerical tasks.

The Hidden Cost of Resampling: How Imbalance Correction Degrades Probability Calibration in Tree Ensembles

TutorialsDGX agent

arXiv:2606.29720v1 Announce Type: cross Abstract: Resampling methods such as SMOTE and random under/over-sampling are standard tools for class-imbalanced classification, almost always evaluated by min

The NTNU System at the S&I Challenge 2025 SLA Open Track

Model ReleasesDGX agent

arXiv:2506.05121v3 Announce Type: replace Abstract: A recent line of research on spoken language assessment (SLA) employs neural models such as BERT and wav2vec 2.0 (W2V) to evaluate speaking proficie

ThinkProbe: Beyond Accuracy -- Structural Profiling of Open-Ended LLM Reasoning Traces via Non-Generative Thought Graphs

ResearchDGX agent

arXiv:2606.29067v1 Announce Type: new Abstract: We present ThinkProbe, a framework for structural analysis of LLM reasoning traces. ThinkProbe converts each trace into a Thought Graph a directed graph

Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding

Model ReleasesDGX agent

arXiv:2601.04693v2 Announce Type: replace Abstract: Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scar

Timesteps of Mamba Align with Human Reading Times

SafetyDGX agent

arXiv:2606.29904v1 Announce Type: new Abstract: This study demonstrates an alignment of per-word processing time in a popular state-space language model Mamba and human readers. In Mamba, the recurren

Towards Physical Intuitions for Alignment Dynamics: A Case Study With Randomness Crystallization

SafetyDGX agent

arXiv:2606.29933v1 Announce Type: new Abstract: The alignment of language models is typically studied through the lens of capability benchmarks, but the dynamics of how models change during post-train

← Previous
1…3031323334…129
Next →