AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
7 Jul 2026

LLM-based Human Simulations Have Not Yet Been Reliable

SafetyDGX agent

arXiv:2501.08579v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current

LLMs Encode Harmfulness and Refusal Separately

Model ReleasesDGX agent

arXiv:2507.11878v5 Announce Type: replace Abstract: LLMs are trained to refuse harmful instructions, but do they truly understand harmfulness beyond just refusing? Prior work has shown that LLMs' refu

LP-SFT: Local-Preserving Supervised Fine-Tuning via Multimodal Entropy Structure

Local AiDGX agent

arXiv:2607.04733v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) is the standard approach for adapting pretrained language models to downstream domains, yet it often improves target-domain


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

Model ReleasesDGX agent

arXiv:2607.02763v1 Announce Type: new Abstract: Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recorded speech, limiting the reach of speech-LLM metho

Measuring and Mitigating Post-hoc Rationalization in Reverse Chain-of-Thought Generation

ResearchDGX agent

arXiv:2602.14469v3 Announce Type: replace Abstract: Reverse Chain-of-Thought Generation (RCG) synthesizes reasoning traces from query-answer pairs, but it risks producing post-hoc rationalizations: wh

Mechanism-level routing failure in LLMs over Lean-verified algebraic structures

Model ReleasesDGX agent

arXiv:2607.04534v1 Announce Type: new Abstract: We present an empirical study of structural routing failure in large language models (LLMs) over a formally verified algebraic corpus. The task requires

Memory-Efficient FastText: A Comprehensive Approach Using Double-Array Trie Structures and Mark-Compact Memory Management

Model ReleasesDGX agent

arXiv:2506.01254v2 Announce Type: replace Abstract: FastText remains a practical choice for industrial word representation because it can synthesize vectors for out-of-vocabulary words from character

Memory-Orchestrated Semantic System (MOSS): An Auditable Agentic Memory Architecture

Local AiDGX agent

arXiv:2607.04391v1 Announce Type: new Abstract: Long-term memory remains a structural weakness of AI agents. The dominant approach, retrieval-augmented generation (RAG), relies on embedding-based simi

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

ResearchDGX agent

arXiv:2607.03540v1 Announce Type: new Abstract: Detecting mental health disorders in a timely manner is an important societal challenge. NLP and machine learning (ML) methods used to assist with detec

MIRAGE: Defending Long-Form RAG Against Misinformation Pollution

ApplicationsDGX agent

arXiv:2607.05069v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) improves factuality by grounding LLMs in external evidence, but real-world retrieval is often polluted: semanticall

MORE: A Multilingual Document Parsing Benchmark and Evaluation

Model ReleasesDGX agent

arXiv:2607.02956v1 Announce Type: cross Abstract: Multilingual documents encapsulate rich regional cultures, scientific discoveries, and historical records. Parsing this content into structured, machi

MTEB-PT: A Text Embedding Benchmark for Brazilian Portuguese

Model ReleasesDGX agent

arXiv:2607.04581v1 Announce Type: new Abstract: Text embeddings for Portuguese have no dedicated benchmark: evaluation rests on translated corpora such as English MS MARCO or on thin multilingual cove

Multi-Large Language Model Orchestrated Severity Assessment of Clinical Records (MOSAIC)

Local AiDGX agent

arXiv:2607.05032v1 Announce Type: new Abstract: Background: Disease severity is a multidimensional construct difficult to capture with rule-based approaches in Electronic Healthcare Records (EHR). Age

OpenSIR: Open-Ended Self-Improving Reasoner

ResearchDGX agent

arXiv:2511.00602v4 Announce Type: replace Abstract: Recent advances in large language model (LLM) reasoning through reinforcement learning rely on annotated datasets for verifiable rewards, which may

Optimizing Large Language Models for Causality Assessment in Pharmacovigilance: Developing a Performance Metric as Objective for Bayesian Hyperparameter Optimization

Model ReleasesDGX agent

arXiv:2607.03704v1 Announce Type: new Abstract: Background: Growing individual case safety report (ICSR) volumes have intensified demand for scalable automated causality assessment. Large Language Mod

Ossetic-COT: Designing a morphologically annotated corpus and morphological analyzer for Ossetic

ResearchDGX agent

arXiv:2607.04895v1 Announce Type: new Abstract: In this work we present the first morphologically annotated corpus for Iron Ossetic that conforms to the Universal Dependencies schema. The corpus inclu

PAST-TIDE: Prototype-Anchored Statement Tuning with Topic-Invariant Normalization for Stance Detection

ResearchDGX agent

arXiv:2607.04690v1 Announce Type: new Abstract: We introduce PAST-TIDE, our stance detection system addressing both subtasks of the StanceNakba Shared Task at NakbaNLP@LREC-COLING 2026. The main idea

Pathways of Visual Information Flow in Vision-Language Models

ResearchDGX agent

arXiv:2607.03358v1 Announce Type: cross Abstract: We study how visual information is routed in vision-language models (VLMs). Using causal patching on controlled synthetic and natural datasets, we fin

PraMem: Practice-derived Experiential Memory for Long-horizon Behavior Prediction

ResearchDGX agent

arXiv:2607.02881v1 Announce Type: new Abstract: Long-horizon behavior prediction aims to infer a user's next action based on a lengthy historical sequence, playing a crucial role in artificial intelli

Predicting the Emergence of Induction Heads in Language Model Pretraining

ResearchDGX agent

arXiv:2511.16893v3 Announce Type: replace Abstract: Specialized attention heads dubbed induction heads (IHs) have been argued to underlie the remarkable in-context learning capabilities of modern lang

ProACT: Towards Breakdown-Aware Proactive Agent in Multi-User Collaboration

Model ReleasesDGX agent

arXiv:2607.03730v1 Announce Type: new Abstract: Conversational agents are increasingly embedded in human collaborative work, yet they remain fundamentally passive and reactive: they respond to explici

Progressive Disclosure for LLM-Maintained Wiki Knowledge Bases: a Preregistered Ablation

AgentsDGX agent

arXiv:2607.04576v1 Announce Type: new Abstract: LLM agents increasingly answer questions against knowledge bases they help maintain. A common intuition holds that progressive disclosure, a compact cat

Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR

ResearchDGX agent

arXiv:2607.05224v1 Announce Type: new Abstract: Code-switching (CS), alternating languages within the same utterance, poses significant challenges for automatic speech recognition (ASR) due to limited

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space

ResearchDGX agent

arXiv:2607.02907v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have achieved remarkable progress but still struggle with complex visual reasoning tasks requiring multi-step

psytechlab at CLPsych 2026: Utilising Natural Language Processing methods and Large Language Models for Social Media Text Analysis

ResearchDGX agent

arXiv:2607.03003v1 Announce Type: new Abstract: Social media posts are a rich and valuable source of data for analyzing mental health states and users' well-being using automated analysis tools. In th

RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain

Model ReleasesDGX agent

arXiv:2607.05171v1 Announce Type: new Abstract: Language understanding in the brain is context-dependent, varying across experimental stimuli and individuals, which makes it difficult to build computa

Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance

ResearchDGX agent

arXiv:2607.05113v1 Announce Type: new Abstract: Imagine two users interact with the same LLM. One has been told it is the cutting-edge flagship model; the other, an older, weaker model. They walk away

Re:Form -- Reducing Human Annotations in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

Model ReleasesDGX agent

arXiv:2507.16331v4 Announce Type: replace Abstract: Existing informal language-based (e.g., human language) Large Language Models (LLMs) trained with Reinforcement Learning (RL) face a significant cha

Reinforcement Learning for Data-Efficient Code-Switched ASR

SafetyDGX agent

arXiv:2607.02757v1 Announce Type: new Abstract: Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language bound

Rethinking AI-Generated Text Detection: A Strong Baseline and the Distribution-Shift Problem That Remains

Model ReleasesDGX agent

arXiv:2607.03680v1 Announce Type: cross Abstract: Recent AI-generated text detection work often introduces a new benchmark together with a specialized detector tailored to it. We revisit this practice

Rethinking Scientific Discovery in an Agentic Era

AgentsDGX agent

arXiv:2607.03863v1 Announce Type: new Abstract: Artificial intelligence has advanced scientific discovery, but most AI4Science systems remain fragmented tools that rely on humans to coordinate problem

S-DiverSe: Spanish Diverse Speech

ResearchDGX agent

arXiv:2607.03207v1 Announce Type: new Abstract: Automatic speech recognition (ASR) has advanced remarkably for standard speech, yet speech affected by neurological conditions remains a challenge. We p

SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis

Model ReleasesDGX agent

arXiv:2607.05259v1 Announce Type: new Abstract: Sentiment analysis has been a primary domain under Natural Language Processing (NLP) from its inception as it plays a vital role in both real-world and

SelfMem: Self-Optimizing Memory for AI Agents

AgentsDGX agent

arXiv:2607.03726v1 Announce Type: new Abstract: While current AI agents support increasingly long context windows, tool use, and skill execution for long-horizon tasks, they still require memory syste

Semantic Homogenization in Italian Popular Music: A Diachronic Analysis

ResearchDGX agent

arXiv:2607.04832v1 Announce Type: new Abstract: In recent years, studies have revealed a decline in semantic variety across popular music lyrics, particularly in English-language songs on streaming pl

Semantic Integration and Lexical Expectation Shape N400 and P600 Dynamics During Naturalistic Reading

Local AiDGX agent

arXiv:2607.04107v1 Announce Type: new Abstract: Word surprisal is a well-established computational predictor of human neural responses during language comprehension, but it remains less clear whether

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning

Model ReleasesDGX agent

arXiv:2603.23483v2 Announce Type: replace-cross Abstract: Agentic multimodal large language models (MLLMs) (e.g., OpenAI o3 and Gemini Agentic Vision) achieve remarkable reasoning capabilities through

Spinning Straw into Gold: Relabeling LLM Agent Trajectories in Hindsight for Successful Demonstrations

AgentsDGX agent

arXiv:2607.04235v1 Announce Type: new Abstract: Large language model agents operate in partially observable, long-horizon settings where obtaining supervision remains a major bottleneck. We address th

Streaming Neural Speech Codecs through Time-Invariant Representations

ResearchDGX agent

arXiv:2607.05250v1 Announce Type: new Abstract: Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized rep

TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding

Model ReleasesDGX agent

arXiv:2607.03236v1 Announce Type: new Abstract: Diffusion language models (DLLMs) generate text by iteratively denoising masked positions, exposing a trajectory of predictive distributions rather than

TAMA: A Human-AI Collaborative Thematic Analysis Framework Using Multi-Agent LLMs for Clinical Interviews

AgentsDGX agent

arXiv:2503.20666v2 Announce Type: replace-cross Abstract: Thematic analysis (TA) is a widely used qualitative approach for uncovering latent meanings in unstructured text data. TA provides valuable in

Teaching Code LLMs to Reason with Intermediate Formal Specifications

Model ReleasesDGX agent

arXiv:2607.04232v1 Announce Type: cross Abstract: Unlike natural-language specifications, executable formal specifications provide machine-checkable constraints for verifying, debugging, and repairing

The Classics at SemEval-2026 Task 3: Combining Transformer Models and LLM-Generated Annotations for Dimensional Aspect-Based Sentiment Analysis

ResearchDGX agent

arXiv:2607.03414v1 Announce Type: new Abstract: This paper presents an approach to the SemEval-2026 Task 3: Dimensional Aspect-Based Sentiment Analysis. We investigate methods for moving beyond tradit

The syntax of wh-agreement in Yemeni Ibbi Arabic

ResearchDGX agent

arXiv:2607.04986v1 Announce Type: new Abstract: This article tackles an important phenomenon in the syntax of Yemeni Ibbi Arabic (YIA), viz., wh-agreement, a phenomenon common to several languages inc

The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices

Model ReleasesDGX agent

arXiv:2603.18482v2 Announce Type: replace Abstract: Standard decoding strategies for text generation, including top-k, nucleus sampling, and contrastive search, select tokens based on likelihood, rest

Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

Model ReleasesDGX agent

arXiv:2602.13517v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive reasoning capabilities by scaling test-time compute via long Chain-of-Thought (CoT). Howev

Token-level Response-visual Attention Guidance for Multimodal LLMs Knowledge Distillation

ResearchDGX agent

arXiv:2607.02593v1 Announce Type: cross Abstract: While knowledge distillation (KD) is widely adopted for training lightweight models by leveraging supervision from larger teacher models, relying sole

TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior

Model ReleasesDGX agent

arXiv:2512.20757v2 Announce Type: replace Abstract: Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs). Despite the importance of tokeniz

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language

ResearchDGX agent

arXiv:2607.04515v1 Announce Type: new Abstract: Efik, a tonal language spoken by about 3 million second language speakers and 1.5 million native speakers in Southeastern Nigeria, remains underrepresen

TRACER: Early Failure Detection for Task-Oriented Dialogue

ResearchDGX agent

arXiv:2607.03974v1 Announce Type: new Abstract: Task-oriented dialogue systems often fail before the final breakdown is obvious, but most evaluation only measures failure after the conversation has al

Train Smarter, Not Longer: Memorization-Guided Data Reuse for Efficient LLM Training

TutorialsDGX agent

arXiv:2607.04969v1 Announce Type: cross Abstract: The training paradigm of large language models has shifted from traditional one-pass training to multi-epoch training, as reasonable reuse of limited

TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking

Model ReleasesDGX agent

arXiv:2410.15135v5 Announce Type: replace Abstract: With the surge of online misinformation, Large Language Models (LLMs) and Reasoning Large Language Models (RLMs) serving as Automatic Fact-Checking

Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees

SafetyDGX agent

arXiv:2607.04430v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in question answering (QA) systems, yet they may generate hallucinated or misaligned responses wi

Variable Bit-width Quantization: Learning Per-Group Precision for 'Bigger-but-Smaller' Language Models

Model ReleasesDGX agent

arXiv:2607.02893v1 Announce Type: cross Abstract: Low-bit quantization shrinks language models but treats precision as a single global hyper-parameter: every weight uses the same bit-width. We introdu

VCB Bench: An Evaluation Benchmark for Audio-Grounded Large Language Model Conversational Agents

Model ReleasesDGX agent

arXiv:2510.11098v5 Announce Type: replace-cross Abstract: Recent advances in large audio language models (LALMs) have greatly enhanced multimodal conversational systems. However, existing benchmarks r

What You See Is What You Get: Observation-Aligned Supervision for Chart-to-Code Generation

ResearchDGX agent

arXiv:2607.04726v1 Announce Type: new Abstract: Chart-to-code generation is commonly trained with supervised fine-tuning on reference plotting scripts, implicitly treating the gold code as a fully obs

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

SafetyDGX agent

arXiv:2607.05132v1 Announce Type: cross Abstract: As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents tha

When Users Are Happy but Agents Are Wrong: Multi-Dimensional Evaluation of Tool-Augmented Dialogue

Model ReleasesDGX agent

arXiv:2510.19186v3 Announce Type: replace Abstract: Evaluating conversational AI systems that use external tools is challenging, as errors can arise from complex interactions among user, agent, and to

When Words Predict Workload

Local AiDGX agent

arXiv:2607.04951v1 Announce Type: cross Abstract: Standard distributed ac{llm} schedulers rely on static token counts or rolling latency averages, making them susceptible to failures on statutorily co

Who's Behind It? Annotating and Extracting Conspiratorial Actors from German Telegram Posts

ResearchDGX agent

arXiv:2607.04962v1 Announce Type: new Abstract: Conspiracy theories commonly attribute important events to the actions of powerful and secretive actors. While computational research has largely focuse

← Previous
1…2526272829…129
Next →