AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
27 Jul 2026

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

Model ReleasesDGX agent

arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contaminati

Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing

TutorialsDGX agent

arXiv:2607.21758v1 Announce Type: new Abstract: Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that produc

InteractComp: Evaluating Search Agents With Ambiguous Queries

Model ReleasesDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.24668v2 Announce Type: replace Abstract: Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that us

Interpretable Depression Detection from Social Media Text Using LLM-Derived Embeddings

ResearchDGX agent

arXiv:2506.06616v2 Announce Type: replace Abstract: Accurate and interpretable detection of depressive language in social media can support early identification of mental health conditions and inform

J-CoT: Chain-of-Thought in J-Space

Model ReleasesDGX agent

arXiv:2607.21981v1 Announce Type: new Abstract: Chain-of-thought prompting improves language-model reasoning by carrying intermediate states across successive computation steps. However, relying on na

Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms

Model ReleasesDGX agent

arXiv:2607.21780v1 Announce Type: new Abstract: Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into the

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

Model ReleasesDGX agent

arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficu

LeAct: Learning to Reason from Expert Actions

Model ReleasesDGX agent

arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely un

Learning to Reason for Factuality

TutorialsDGX agent

arXiv:2508.05618v2 Announce Type: replace Abstract: Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substant

Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets

Model ReleasesDGX agent

arXiv:2607.21692v1 Announce Type: cross Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models

Local AiDGX agent

arXiv:2607.21936v1 Announce Type: new Abstract: Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing re

LMEB: Long-horizon Memory Embedding Benchmark

Model ReleasesDGX agent

arXiv:2603.12572v5 Announce Type: replace Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchm

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers

SafetyDGX agent

arXiv:2607.15456v2 Announce Type: replace-cross Abstract: Looped, weight-tied Transformers reduce parameters by reusing a single block, but decoding still stores a separate K/V cache for every recurre

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

AgentsDGX agent

arXiv:2508.12393v3 Announce Type: replace Abstract: The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet curr

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

ResearchDGX agent

arXiv:2607.22100v1 Announce Type: new Abstract: Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into t

MoE^2-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation

Model ReleasesDGX agent

arXiv:2607.21978v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures have been widely adopted in large language models, yet parameter-efficient fine-tuning (PEFT) for MoE models rema

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

SafetyDGX agent

arXiv:2607.21653v1 Announce Type: cross Abstract: Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstrea

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

SafetyDGX agent

arXiv:2607.22083v1 Announce Type: cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-

On Improving Faithfulness of Podcasts from Documents

ResearchDGX agent

arXiv:2607.21961v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate long-form conversational content such as podcasts from textual sources. While these syste

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

Model ReleasesDGX agent

arXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor

Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

SafetyDGX agent

arXiv:2607.21774v1 Announce Type: new Abstract: Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot stud

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

ResearchDGX agent

arXiv:2607.21655v1 Announce Type: cross Abstract: Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is co

Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster

SafetyDGX agent

arXiv:2607.18292v2 Announce Type: replace-cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account --

Scaling Native Multimodal Pre-Training From Scratch

ResearchDGX agent

arXiv:2607.22043v1 Announce Type: new Abstract: Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of th

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning

SafetyDGX agent

arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

Model ReleasesDGX agent

arXiv:2607.22529v1 Announce Type: new Abstract: LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fund

Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation

ResearchDGX agent

arXiv:2607.22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor li

SwiftMem: Fast Agentic Memory via Query-aware Indexing

Model ReleasesDGX agent

arXiv:2601.08160v2 Announce Type: replace Abstract: Agentic memory systems have become critical for enabling LLM agents to maintain long-term context and retrieve relevant information efficiently. How

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

ResearchDGX agent

arXiv:2607.21971v1 Announce Type: cross Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We h

The Hard Decision Layer: Evidence for Committed Inference in Transformers

Model ReleasesDGX agent

arXiv:2607.21613v1 Announce Type: cross Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Deci

Token-Operations-Oriented Inference Optimization Techniques for Large Models

ApplicationsDGX agent

arXiv:2606.20295v2 Announce Type: replace-cross Abstract: Large model inference optimization serves as a key foundation for supporting the scalable, low-cost, and highly stable operation of large mode

Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents

AgentsDGX agent

arXiv:2607.21887v1 Announce Type: cross Abstract: Foreign language anxiety (FLA) can be a major barrier to second language acquisition (SLA), especially in conversational contexts. With the proliferat

What Matters When Building Universal Multilingual Named Entity Recognition Models?

ResearchDGX agent

arXiv:2601.06347v2 Announce Type: replace Abstract: Recent progress in universal multilingual named entity recognition (NER) has been driven by multilingual transformer models, task-specific architect

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

Model ReleasesDGX agent

arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Wom

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

SafetyDGX agent

arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical align

Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

SafetyDGX agent

arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, rais

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

Model ReleasesDGX agent

arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose m

24 Jul 2026

A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour

ResearchDGX agent

arXiv:2607.18317v2 Announce Type: replace-cross Abstract: We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open

A Unified Moral-Value Dataset for Instruction Tuning

SafetyDGX agent

arXiv:2607.21279v1 Announce Type: new Abstract: Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human

Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs

SafetyDGX agent

arXiv:2607.21291v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost. Existing

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

ResearchDGX agent

arXiv:2607.21424v1 Announce Type: new Abstract: Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disen

Anti-Periodic Positional Encoding: Mobius Boundary Conditions Make In-Context Retrieval Reliable

ResearchDGX agent

arXiv:2607.21405v1 Announce Type: new Abstract: Mobius RoPE is a rotary positional encoding built on the anti-periodic frequency ladder heta_i=pi(2i+1)/N: every rotation plane advances by an odd multi

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

Model ReleasesDGX agent

arXiv:2607.20596v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE fami

AsymVerify at SemEval-2026 Task 6: Asymmetric Confidence-Gated Verification for Political Evasion Detection

ResearchDGX agent

arXiv:2607.20439v1 Announce Type: new Abstract: Political evasion is difficult to detect because evasive answers often appear cooperative while avoiding concrete commitment. We present AsymVerify, a c

Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets

SafetyDGX agent

arXiv:2607.20441v1 Announce Type: new Abstract: Every information ecosystem produces beliefs that shape strategic decisions. Both human analysts and AI systems inherit the blind spots of their informa

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

Model ReleasesDGX agent

arXiv:2607.21340v1 Announce Type: new Abstract: In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defens

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

ResearchDGX agent

arXiv:2607.20935v1 Announce Type: cross Abstract: Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning mo

Content Anonymization for Privacy in Long-form Audio

ResearchDGX agent

arXiv:2510.12780v3 Announce Type: replace-cross Abstract: Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmar

CSPF: A Constrained Shared-Private Fusion Method for Non-Verifiable Preference Evaluation

SafetyDGX agent

arXiv:2607.20862v1 Announce Type: new Abstract: At present, reliable evaluation of non-verifiable tasks remains challenging. Existing approaches often fail to adequately capture the diverse evaluative

CultureTalk-ID: A Multi-Task Dialogue Benchmark for Cultural Commonsense in Indonesian Local Languages

Model ReleasesDGX agent

arXiv:2607.21016v1 Announce Type: new Abstract: Culture is lived through conversation, yet existing Indonesian cultural commonsense benchmarks evaluate LLMs on short and isolated prompts, stripping aw

DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

Model ReleasesDGX agent

arXiv:2607.20465v1 Announce Type: cross Abstract: The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how

DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

Model ReleasesDGX agent

arXiv:2603.11838v2 Announce Type: replace Abstract: Large language models pretrained on internet-scale data risk lookahead bias in forecasting tasks, as they may have already seen the true outcome dur

Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content

ResearchDGX agent

arXiv:2607.20446v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used by students to generate natural language responses and program code, there is growing interest in

Domyn-Small: A European 10B Reasoning Language Model

Model ReleasesDGX agent

arXiv:2607.20448v1 Announce Type: new Abstract: We introduce Domyn-Small, a 10-billion-parameter open-weight reasoning language model released under the MIT license. Domyn-Small is the product of an i

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages

Model ReleasesDGX agent

arXiv:2607.21540v1 Announce Type: new Abstract: We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2.

Evaluating the Effectiveness of Persona Simulation in Opinion Prediction with GPT-4.1

Model ReleasesDGX agent

arXiv:2607.20589v1 Announce Type: new Abstract: Persona simulation involves utilizing large language models (LLMs) to anticipate human choices or interactions based on specific characteristic informat

From a Word-Level Dictionary to Sentence-Level Semantics: Multilingual Grievance Labelling with Contextual Models

Model ReleasesDGX agent

arXiv:2607.20946v1 Announce Type: new Abstract: Grievance is one of the warning signs analysts look for when assessing threats of violence. It is increasingly measured at scale from online text, most

GaugeQuant: Online Learning of Quantization-Optimal Bases from LLM Symmetries

Model ReleasesDGX agent

arXiv:2607.20757v1 Announce Type: cross Abstract: Transformers are known to have internal continuous symmetries that leave outputs invariant, while modifying quantization. GaugeQuant leverages this in

GLAN-QnA-KR: A Seedless Taxonomy-Driven Korean Instruction Corpus

Model ReleasesDGX agent

arXiv:2607.20443v1 Announce Type: new Abstract: We release GLAN-QnA-KR, a 303,581-row openly redistributable Korean instruction-QA corpus produced via the seedless taxonomy-driven GLAN synthesis pipel

Gumbel Distillation for Parallel Text Generation

TutorialsDGX agent

arXiv:2603.22216v2 Announce Type: replace Abstract: The slow, sequential nature of autoregressive (AR) language models has driven the adoption of parallel decoding methods. However, these non-AR model

← Previous
1…1920212223…129
Next →