AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
25 Jun 2026

IndicContextEval: A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages

Model ReleasesDGX agent

arXiv:2606.19157v2 Announce Type: replace-cross Abstract: AudioLLMs enable speech recognition conditioned on textual prompts such as domain descriptions or entity lists. However, it remains unclear wh

Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations

ResearchDGX agent

arXiv:2606.25383v1 Announce Type: new Abstract: As previous research on annotator disagreement in discourse phenomena has shown, understanding text coherence varies considerably from one individual to

Invisible to humans, visible to machines: a preregistered audit of Unicode fidelity across four biomedical bibliographic APIs


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Research
DGX agent

arXiv:2606.24897v1 Announce Type: cross Abstract: Biomedical text mining, scientometrics, and the construction of training corpora for biomedical large language models (LLMs) all assume that the abstr

Is GraphRAG Needed? From Basic RAG to Graph-/Agentic Solutions with Context Optimization

AgentsDGX agent

arXiv:2606.25656v1 Announce Type: new Abstract: As advanced RAG variants like GraphRAG and Agentic RAG emerge, one leading question is when and how to use them. Here, we introduce a framework for diff

JetSpec: Breaking the Scaling Ceiling of Speculative Decoding with Parallel Tree Drafting

HardwareDGX agent

arXiv:2606.18394v2 Announce Type: replace Abstract: Speculative decoding (SD) accelerates autoregressive Large Language Models (LLMs) by drafting multiple tokens and verifying them in parallel, but it

Learning Diachronic Representations of Ancient Greek Letterforms

ApplicationsDGX agent

arXiv:2606.24984v1 Announce Type: cross Abstract: Learning representations that remain robust across centuries of variation in handwriting is a key challenge in diachronic representation learning. Tak

Learning task-specific subspaces via interventional post-training of speech foundation models

TutorialsDGX agent

arXiv:2606.17967v2 Announce Type: replace Abstract: Speech foundation models, pre-trained on large corpora of unlabelled speech data, produce general-purpose representations which are useful across ta

Learning to Erase Private Knowledge from Multi-Documents for Retrieval-Augmented Large Language Models

ResearchDGX agent

arXiv:2504.09910v2 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) is a promising technique for applying LLMs to proprietary domains. However, retrieved documents may contain sen

LLM-ACES: Closed-Loop Discovery of Dynamical Systems with LLM-Guided Adaptive Search

ResearchDGX agent

arXiv:2606.25039v1 Announce Type: cross Abstract: Recovering governing Ordinary Differential Equations (ODEs) from data is a central challenge in modeling dynamical systems across scientific domains.

LLM-Based Scientific Peer Review: Methods, Benchmarks, and Reliability Challenges

SafetyDGX agent

arXiv:2606.25057v1 Announce Type: new Abstract: The rapid growth of scientific submissions has pushed traditional peer review toward its scalability limits, motivating the exploration of large languag

LLM Performance on a Real, Double-Marked GCSE Benchmark

Model ReleasesDGX agent

arXiv:2606.24973v1 Announce Type: new Abstract: We introduce a dataset of 32,534 double-marked real student responses to GCSE mock exams (GCSEs are the UK's national exams, taken at age ~16), spanning

Measuring Research Difficulty of Academic Papers: A Case Study in Natural Language Processing

ApplicationsDGX agent

arXiv:2606.25307v1 Announce Type: cross Abstract: With the rapid growth of the number of academic papers, systematically evaluating the difficulty of research and its relationship to academic impact o

MedGuards: Multi-Agent System for Reliable Medical Error Detection and Correction

SafetyDGX agent

arXiv:2606.25651v1 Announce Type: new Abstract: As Large Language Models (LLMs) are increasingly deployed in healthcare settings, accurate error detection and correction in generated or existing text

MedLayBench-V: A Large-Scale Benchmark for Expert-Lay Semantic Alignment in Medical Vision Language Models

Model ReleasesDGX agent

arXiv:2604.05738v2 Announce Type: replace Abstract: Medical Vision-Language Models (Med-VLMs) have achieved expert-level proficiency in interpreting diagnostic imaging. However, current models are pre

Membox: Weaving Topic Continuity into Long-Range Memory for LLM Agents

AgentsDGX agent

arXiv:2601.03785v3 Announce Type: replace Abstract: Long-term human-agent dialogues are organized by topic continuity: adjacent turns often develop the same goal, plan, problem, or event, while relate

Memory Makes the Difference: Evaluating How Different Memory Roles Shape Conversational Agents

AgentsDGX agent

arXiv:2606.25361v1 Announce Type: new Abstract: Prior research on memory mechanism in RAG-based conversational system has emphasized how memory is stored and retrieved. However, far less is known abou

Multilingual Hematology Visual Question Answering Dataset

Model ReleasesDGX agent

arXiv:2606.25246v1 Announce Type: cross Abstract: Vision Language Models (VLMs) have shown promising capabilities in medical image analysis by jointly understanding visual and textual information for

Narrative Feature or Structured Feature? A Study of Large Language Models to Identify Cancer Patients at Risk of Heart Failure

SafetyDGX agent

arXiv:2403.11425v4 Announce Type: replace-cross Abstract: Cancer treatments are known to introduce cardiotoxicity, negatively impacting outcomes and survivorship. Identifying cancer patients at risk o

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

Model ReleasesDGX agent

arXiv:2606.26050v1 Announce Type: cross Abstract: Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ('Sue cried because'), it r

Neural Machine Translation for Low-Resource Tangkhul--English

SafetyDGX agent

arXiv:2606.25365v1 Announce Type: new Abstract: We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair. Tangkhul is a severely under-resourced Tibeto-Bu

Neural Scaling Universality: If Exponents Are Fixed, Time to Understand Coefficients

ResearchDGX agent

arXiv:2606.25008v1 Announce Type: cross Abstract: Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute. This position paper argues that t

OPERA: Aligning Open-Ended Reasoning via Objective Perplexity-based Reinforcement Learning

SafetyDGX agent

arXiv:2606.25757v1 Announce Type: new Abstract: Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open

Optimizing Abstractive Summarization With Fine-Tuned PEGASUS

ResearchDGX agent

arXiv:2606.25462v1 Announce Type: new Abstract: Abstractive text summarization is the technique of generating a short and concise summary comprising the salient ideas of a source text without making a

Overview of HIPE-2026: Person-Place Relation Extraction from Multilingual Historical Texts

ResearchDGX agent

arXiv:2606.25935v1 Announce Type: new Abstract: Was this person ever at that place, and if so, when? Answering such questions from noisy, multilingual historical documents is the central challenge of

Paid Voices vs. Public Feeds: Interpretable Cross-Platform Theme-Based Analysis of Climate Discourse

SafetyDGX agent

arXiv:2601.13317v2 Announce Type: replace Abstract: Climate discourse online shapes public understanding of climate change and informs political and policy debate, yet it unfolds across structurally d

Perfect Detection, Failed Control: The Geometry of Knowing vs. Steering in Language Models

Model ReleasesDGX agent

arXiv:2606.24952v1 Announce Type: new Abstract: A central aspiration of mechanistic interpretability is controllability: if we know where a behavior is represented in a model's activations, we should

PhoneBuddy: Training Open Models for Agentic Phone Use

AgentsDGX agent

arXiv:2606.23049v2 Announce Type: replace Abstract: Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult bec

PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models

Model ReleasesDGX agent

arXiv:2606.25442v1 Announce Type: new Abstract: Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. Ho

Position: Reasoning After Perception Means Reasoning Without Vision

ResearchDGX agent

arXiv:2507.16863v2 Announce Type: replace-cross Abstract: A common belief in multimodal research is that the perceptual weaknesses of vision--language models can be compensated by stronger language re

Privacy-Aware Visual Language Models

Model ReleasesDGX agent

arXiv:2405.17423v4 Announce Type: replace-cross Abstract: As Visual Language Models (VLMs) become increasingly embedded in everyday applications, ensuring they can recognise and appropriately handle p

Probing in the Wild: A Case Study of Self-Supervised Speech Representations on Mandarin Sub-dialects with Unsupervised Articulatory Analysis

ApplicationsDGX agent

arXiv:2606.25459v1 Announce Type: new Abstract: While self-supervised speech models have achieved strong performance across speech tasks, relatively little is known about how their internal phonetic r

RAS: Measuring LLM Safety Through Refusal Alignment

Model ReleasesDGX agent

arXiv:2606.25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their

RAVEN: Long-Horizon Reasoning & Navigation with a Visuo-Spatio-Temporal Memory

AgentsDGX agent

arXiv:2606.25206v1 Announce Type: cross Abstract: Long-term robot deployment requires a compact and scalable memory that preserves fine-grained visual semantics, grounds observations in space and time

Real-Time Voice AI Hears but Does Not Listen

Model ReleasesDGX agent

arXiv:2606.26083v1 Announce Type: new Abstract: Speech conveys information through both words and vocal delivery. We evaluate four leading production realtime voice systems-OpenAI's GPT Realtime 2, Go

Reclaim Evaluation: A Lossy Memory Is Worse Than an Empty One

Model ReleasesDGX agent

arXiv:2606.25449v1 Announce Type: new Abstract: A language model's memory can be worse than having no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behind it,

Reinforcement Learning Improves Traversal of Parametric Knowledge in LLMs

ResearchDGX agent

arXiv:2511.05933v2 Announce Type: replace Abstract: Reinforcement learning (RL) is often credited with improving language model reasoning at the expense of knowledge. We challenge this narrative by sh

Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning

ResearchDGX agent

arXiv:2606.25568v1 Announce Type: new Abstract: Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on English-centric training resources and benchmarks

Robustness assessment of large audio language models in multiple-choice evaluation

Model ReleasesDGX agent

arXiv:2510.04584v2 Announce Type: replace Abstract: Recent advances in large audio language models (LALMs) have primarily been assessed using a multiple-choice question answering (MCQA) framework. How

Same Evidence, Different Answer: Auditing Order Sensitivity in Multimodal Large Language Models

Model ReleasesDGX agent

arXiv:2606.26079v1 Announce Type: new Abstract: Standard benchmarks for multimodal large language models (MLLMs) score each item on one canonical ordering and miss whether order-irrelevant shuffling c

SARA: Unlocking Multilingual Knowledge in Mixture-of-Experts via Semantically Anchored Routing Alignment

Model ReleasesDGX agent

arXiv:2606.25821v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) architectures have emerged as an increasingly influential paradigm as they offer a strategic balance between parameter s

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis

Model ReleasesDGX agent

arXiv:2606.25369v1 Announce Type: cross Abstract: While large language model (LLM)-based text-to-speech (TTS) systems have achieved high-quality speech synthesis, most existing systems focus on Englis

Scale or Reason? A Compute-Equivalent Analysis of Reasoning Distillation

ResearchDGX agent

arXiv:2509.22193v2 Announce Type: replace Abstract: Distilling reasoning traces from strong teacher models has become the standard recipe for building capable small language models. Yet reasoning trac

Security and Privacy in Retrieval-Augmented Generation: Architectures, Threats, Defenses, and Future Directions for Building Trustworthy Systems

Local AiDGX agent

arXiv:2606.25533v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) has emerged as a dominant paradigm for enhancing large language models with external knowledge. By coupling retri

SFL-MTSC: Leveraging Semantic Frame-Level Multi-Task Self-Consistency for Robust Multi-Intent Spoken Language Understanding

Model ReleasesDGX agent

arXiv:2606.25552v1 Announce Type: new Abstract: Prompt-based spoken language understanding (SLU) with large language models (LLMs) often suffers from inconsistent intent--slot structures due to decodi

Small edits, large models: How Wikipedia advocacy shapes LLM values

Model ReleasesDGX agent

arXiv:2606.24890v1 Announce Type: new Abstract: Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in near

Space-Efficient Language Generation in the Limit

ResearchDGX agent

arXiv:2606.25777v1 Announce Type: cross Abstract: We initiate a resource-aware theory of extit{language generation in the limit} under the minimal constraint of space efficiency. In our framework, a l

Spam and Sentiment Detection in Arabic Tweets Using MARBERT Model

ResearchDGX agent

arXiv:2606.25495v1 Announce Type: new Abstract: Saudi Telecom Company (STC) is among the most popular companies in Saudi Arabia, with many customers. Yet, there is still a big room for improvement in

SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs

Model ReleasesDGX agent

arXiv:2602.06566v3 Announce Type: replace-cross Abstract: Despite recent successes, test-time scaling -- i.e., dynamically expanding the token budget during inference as needed -- remains brittle for

Speech Codec Probing from Semantic and Phonetic Perspectives

Model ReleasesDGX agent

arXiv:2603.10371v2 Announce Type: replace-cross Abstract: Speech tokenizers are essential for connecting speech to large language models (LLMs) in multimodal systems. Speech tokenizers are expected to

SpeechEQ: Benchmarking Emotional Intelligence Quotient in Socially Aware Voice Conversational Models

Model ReleasesDGX agent

arXiv:2606.25990v1 Announce Type: new Abstract: As multimodal conversational systems increasingly engage in spoken interaction, their ability to navigate paralinguistic social cues has become a critic

Staying In Character: Perspective-Bounded Memory For Book-Based Role-Playing Agents

Model ReleasesDGX agent

arXiv:2606.25632v1 Announce Type: new Abstract: Recent LLM role-playing systems build character agents from novels by extracting characters, scenes, and relations. Yet long-narrative role-playing suff

Story Operators: Decomposing the Original o Sequel Transformation in Embedding Space

Model ReleasesDGX agent

arXiv:2606.25379v1 Announce Type: new Abstract: I treat a book as a point in a sentence-embedding space and a literary transformation as an operation on points. Given an original novel and its sequel,

Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding

ResearchDGX agent

arXiv:2601.17917v3 Announce Type: replace-cross Abstract: Diffusion Large Language Models (dLLMs) offer a compelling paradigm for natural language generation, leveraging parallel decoding and bidirect

SyncLoop: A Multimodal Dual-Loop Framework for Self-Improving Mathematical Reasoning

Model ReleasesDGX agent

arXiv:2507.16518v3 Announce Type: replace-cross Abstract: Recent advances in multimodal large language models (MLLMs) have shown impressive reasoning capabilities. However, further enhancing existing

The cognitive, affective, and behavioral expression of self-stigma among people who use drugs in online substance use communities

ResearchDGX agent

arXiv:2606.25143v1 Announce Type: new Abstract: Objectives: To develop a codebook for self-stigma across cognitive, affective, and behavioral domains, and to estimate the prevalence, co-occurrence, an

The Generalization Spectrum: A Chromatographic Approach to Evaluating Learning Algorithms

ResearchDGX agent

arXiv:2606.25450v1 Announce Type: cross Abstract: Traditional evaluations measure a learning algorithm's final performance on an i.i.d. test set, reducing learning to a single aggregate score. This ap

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems

SafetyDGX agent

arXiv:2606.24937v1 Announce Type: cross Abstract: The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the full stack fr

The Interplay of Harness Design and Post-Training in LLM Agents

ResearchDGX agent

arXiv:2606.25447v1 Announce Type: cross Abstract: Tool-integrated LLM agents are often wrapped within a harness: the scaffolding that determines which tools are exposed, how they are described, and wh

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar

SafetyDGX agent

arXiv:2606.26015v1 Announce Type: new Abstract: Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities a

Three Buddhist Vocabularies: Computational Stylometry of the English Pali Canon across Sutta, Vinaya, and Abhidhamma

Model ReleasesDGX agent

arXiv:2606.25372v1 Announce Type: new Abstract: We present a computational stylometric analysis of the Tipitaka across all three Pitakas in English translation, extending earlier work on the Sutta Pit

← Previous
1…3435363738…129
Next →