AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings

DGX agent

arXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contaminati

model-releasesarxiv-cs-cl
27 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Tutorials

Humanly: A Configurable and Traceable Environment for Human-AI Collaborative Writing

DGX agent

arXiv:2607.21758v1 Announce Type: new Abstract: Teachers, conference chairs, and public readers all judge writing from limited evidence, seeing only a finished document and not the process that produc

tutorialsarxiv-cs-cl
27 Jul 2026
Model Releases

InteractComp: Evaluating Search Agents With Ambiguous Queries

DGX agent

arXiv:2510.24668v2 Announce Type: replace Abstract: Language agents have demonstrated remarkable potential in web search and information retrieval. However, many search-agent benchmarks assume that us

model-releasesarxiv-cs-cl
27 Jul 2026
Research

Interpretable Depression Detection from Social Media Text Using LLM-Derived Embeddings

DGX agent

arXiv:2506.06616v2 Announce Type: replace Abstract: Accurate and interpretable detection of depressive language in social media can support early identification of mental health conditions and inform

researcharxiv-cs-cl
27 Jul 2026
Model Releases

J-CoT: Chain-of-Thought in J-Space

DGX agent

arXiv:2607.21981v1 Announce Type: new Abstract: Chain-of-thought prompting improves language-model reasoning by carrying intermediate states across successive computation steps. However, relying on na

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

Khondo: A Multimodal Benchmark for Document Packet Splitting of Bangla Forms

DGX agent

arXiv:2607.21780v1 Announce Type: new Abstract: Document packets, multiple documents concatenated into a single file, are common in government and administrative workflows, yet splitting them into the

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision

DGX agent

arXiv:2603.07025v2 Announce Type: replace Abstract: Speech Large Language Models (LLMs) that understand and follow instructions in many languages are useful for real-world interaction, but are difficu

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

LeAct: Learning to Reason from Expert Actions

DGX agent

arXiv:2607.21856v1 Announce Type: cross Abstract: Modern reasoning models depend on reasoning data, today sourced from human annotations or distilled from stronger LLMs. However, a rich and largely un

model-releasesarxiv-cs-cl
27 Jul 2026
Tutorials

Learning to Reason for Factuality

DGX agent

arXiv:2508.05618v2 Announce Type: replace Abstract: Reasoning Large Language Models (R-LLMs) have significantly advanced complex reasoning tasks but often struggle with factuality, generating substant

tutorialsarxiv-cs-cl
27 Jul 2026
Model Releases

Learning What Matters: Supervising Sparse Attention Routing with Causal Evidence Sets

DGX agent

arXiv:2607.21692v1 Announce Type: cross Abstract: Sparse attention reduces the cost of long contexts by allowing each query to read only selected parts of the input. These selectors are often trained

model-releasesarxiv-cs-cl
27 Jul 2026
Local Ai

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models

DGX agent

arXiv:2607.21936v1 Announce Type: new Abstract: Historical documents act as invaluable knowledge archives but often suffer from illegibility due to physical deterioration and damage. While existing re

local-aiarxiv-cs-cl
27 Jul 2026
Model Releases

LMEB: Long-horizon Memory Embedding Benchmark

DGX agent

arXiv:2603.12572v5 Announce Type: replace Abstract: Memory embeddings are crucial for memory-augmented systems, such as OpenClaw, but their evaluation is underexplored in current text embedding benchm

model-releasesarxiv-cs-cl
27 Jul 2026
Safety

Looped Latent Attention: Cross-Loop KV Compression for Looped Transformers

DGX agent

arXiv:2607.15456v2 Announce Type: replace-cross Abstract: Looped, weight-tied Transformers reduce parameters by reusing a single block, but decoding still stores a separate K/V cache for every recurre

safetyarxiv-cs-cl
27 Jul 2026
Agents

MedKGent: A Large Language Model Agent Framework for Constructing Temporally Evolving Medical Knowledge Graph

DGX agent

arXiv:2508.12393v3 Announce Type: replace Abstract: The rapid expansion of medical literature challenges the scalable structuring of domain knowledge. Knowledge Graphs (KGs) offer a solution, yet curr

agentsarxiv-cs-cl
27 Jul 2026
Research

MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond

DGX agent

arXiv:2607.22100v1 Announce Type: new Abstract: Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into t

researcharxiv-cs-cl
27 Jul 2026
Model Releases

MoE^2-LoRA: When MoE Models Meet MoE-style Low-Rank Adaptation

DGX agent

arXiv:2607.21978v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) architectures have been widely adopted in large language models, yet parameter-efficient fine-tuning (PEFT) for MoE models rema

model-releasesarxiv-cs-cl
27 Jul 2026
Safety

Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning

DGX agent

arXiv:2607.21653v1 Announce Type: cross Abstract: Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstrea

safetyarxiv-cs-cl
27 Jul 2026
Safety

Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Mode

DGX agent

arXiv:2607.22083v1 Announce Type: cross Abstract: We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-

safetyarxiv-cs-cl
27 Jul 2026
Research

On Improving Faithfulness of Podcasts from Documents

DGX agent

arXiv:2607.21961v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used to generate long-form conversational content such as podcasts from textual sources. While these syste

researcharxiv-cs-cl
27 Jul 2026
Model Releases

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

DGX agent

arXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor

model-releasesarxiv-cs-cl
27 Jul 2026
Safety

Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders

DGX agent

arXiv:2607.21774v1 Announce Type: new Abstract: Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot stud

safetyarxiv-cs-cl
27 Jul 2026
Research

Progress Reward Modeling for Robotic Learning: A Comprehensive Survey

DGX agent

arXiv:2607.21655v1 Announce Type: cross Abstract: Robotic learning takes place in dynamic environments with large behavior spaces. A terminal success signal only tells the robot whether the task is co

researcharxiv-cs-cl
27 Jul 2026
Safety

Reliability Scales Inversely: Bigger Language Models Compound Mistakes Faster

DGX agent

arXiv:2607.18292v2 Announce Type: replace-cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability. The knowledge-gap account --

safetyarxiv-cs-cl
27 Jul 2026
Research

Scaling Native Multimodal Pre-Training From Scratch

DGX agent

arXiv:2607.22043v1 Announce Type: new Abstract: Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of th

researcharxiv-cs-cl
27 Jul 2026
Safety

Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning

DGX agent

arXiv:2507.01551v3 Announce Type: replace-cross Abstract: Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~

safetyarxiv-cs-cl
27 Jul 2026
Model Releases

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills

DGX agent

arXiv:2607.22529v1 Announce Type: new Abstract: LLM training is shifting from manual design and annotation to interaction-driven self-evolution. However, existing self-evolutionary methods face a fund

model-releasesarxiv-cs-cl
27 Jul 2026
Research

Small Vision-Language Models Know When They Are Wrong But Cannot Say So: A Two-Model Study of Stated versus Internal Confidence Under Realistic Image Degradation

DGX agent

arXiv:2607.22034v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly deployed on consumer hardware where input images are degraded by compression, camera shake, and poor li

researcharxiv-cs-cl
27 Jul 2026
Model Releases

SwiftMem: Fast Agentic Memory via Query-aware Indexing

DGX agent

arXiv:2601.08160v2 Announce Type: replace Abstract: Agentic memory systems have become critical for enabling LLM agents to maintain long-term context and retrieve relevant information efficiently. How

model-releasesarxiv-cs-cl
27 Jul 2026
Research

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

DGX agent

arXiv:2607.21971v1 Announce Type: cross Abstract: Test-time scaling through iterative self-evolution with environment feedback, as demonstrated by AlphaEvolve, shows remarkable performance gains. We h

researcharxiv-cs-cl
27 Jul 2026
Model Releases

The Hard Decision Layer: Evidence for Committed Inference in Transformers

DGX agent

arXiv:2607.21613v1 Announce Type: cross Abstract: We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Deci

model-releasesarxiv-cs-cl
27 Jul 2026
Applications

Token-Operations-Oriented Inference Optimization Techniques for Large Models

DGX agent

arXiv:2606.20295v2 Announce Type: replace-cross Abstract: Large model inference optimization serves as a key foundation for supporting the scalable, low-cost, and highly stable operation of large mode

applicationsarxiv-cs-cl
27 Jul 2026
Agents

Towards Reducing Foreign Language Anxiety Using Level-Appropriate Embodied Conversational Agents

DGX agent

arXiv:2607.21887v1 Announce Type: cross Abstract: Foreign language anxiety (FLA) can be a major barrier to second language acquisition (SLA), especially in conversational contexts. With the proliferat

agentsarxiv-cs-cl
27 Jul 2026
Research

What Matters When Building Universal Multilingual Named Entity Recognition Models?

DGX agent

arXiv:2601.06347v2 Announce Type: replace Abstract: Recent progress in universal multilingual named entity recognition (NER) has been driven by multilingual transformer models, task-specific architect

researcharxiv-cs-cl
27 Jul 2026
Model Releases

WHBench: Evaluating Frontier LLMs with Expert-in-the-Loop Validation on Women's Health Topics

DGX agent

arXiv:2604.00024v2 Announce Type: replace Abstract: Large language models are increasingly used for medical guidance, but women's health remains under-evaluated in benchmark design. We present the Wom

model-releasesarxiv-cs-cl
27 Jul 2026
Safety

When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas

DGX agent

arXiv:2505.19212v2 Announce Type: replace Abstract: Recent advances in LLMs have enabled their use in complex agentic roles, involving decision-making with humans or other agents, making ethical align

safetyarxiv-cs-cl
27 Jul 2026
Safety

Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity

DGX agent

arXiv:2607.22218v1 Announce Type: new Abstract: Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, rais

safetyarxiv-cs-cl
27 Jul 2026
Model Releases

Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents

DGX agent

arXiv:2607.22014v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose m

model-releasesarxiv-cs-cl
27 Jul 2026
Research

A Situational Speech Synthesizer for Yoruba: System Design, Phonological Rule Architecture, and Orthographic Extensions for Contour

DGX agent

arXiv:2607.18317v2 Announce Type: replace-cross Abstract: We present TTSYoruba, a rule-based concatenative diphone speech synthesizer for Yoruba, deployed at online as part of the YorubaName.com open

researcharxiv-cs-cl
24 Jul 2026
Safety

A Unified Moral-Value Dataset for Instruction Tuning

DGX agent

arXiv:2607.21279v1 Announce Type: new Abstract: Large language models (LLMs) have developed rapidly and become valuable tools in everyday life. However, how to align LLMs to a particular set of human

safetyarxiv-cs-cl
24 Jul 2026
Safety

Adaptive Depth Sparse Framework: Similarity-Driven Resource Allocation for Pre-Trained LLMs

DGX agent

arXiv:2607.21291v1 Announce Type: new Abstract: Large language models (LLMs) achieve strong generation and reasoning performance, but the Transformer architecture incurs high inference cost. Existing

safetyarxiv-cs-cl
24 Jul 2026
Research

An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations

DGX agent

arXiv:2607.21424v1 Announce Type: new Abstract: Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disen

researcharxiv-cs-cl
24 Jul 2026
Research

Anti-Periodic Positional Encoding: Mobius Boundary Conditions Make In-Context Retrieval Reliable

DGX agent

arXiv:2607.21405v1 Announce Type: new Abstract: Mobius RoPE is a rotary positional encoding built on the anti-periodic frequency ladder heta_i=pi(2i+1)/N: every rotation plane advances by an odd multi

researcharxiv-cs-cl
24 Jul 2026
Model Releases

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects

DGX agent

arXiv:2607.20596v1 Announce Type: cross Abstract: Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE fami

model-releasesarxiv-cs-cl
24 Jul 2026
Research

AsymVerify at SemEval-2026 Task 6: Asymmetric Confidence-Gated Verification for Political Evasion Detection

DGX agent

arXiv:2607.20439v1 Announce Type: new Abstract: Political evasion is difficult to detect because evasive answers often appear cooperative while avoiding concrete commitment. We present AsymVerify, a c

researcharxiv-cs-cl
24 Jul 2026
Safety

Belief Propagation in LLM World Models: Measuring Strategic Information Bias with Prediction Markets

DGX agent

arXiv:2607.20441v1 Announce Type: new Abstract: Every information ecosystem produces beliefs that shape strategic decisions. Both human analysts and AI systems inherit the blind spots of their informa

safetyarxiv-cs-cl
24 Jul 2026
Model Releases

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

DGX agent

arXiv:2607.21340v1 Announce Type: new Abstract: In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defens

model-releasesarxiv-cs-cl
24 Jul 2026
Research

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

DGX agent

arXiv:2607.20935v1 Announce Type: cross Abstract: Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning mo

researcharxiv-cs-cl
24 Jul 2026
Research

Content Anonymization for Privacy in Long-form Audio

DGX agent

arXiv:2510.12780v3 Announce Type: replace-cross Abstract: Voice anonymization techniques have been found to successfully obscure a speaker's acoustic identity in short, isolated utterances in benchmar

researcharxiv-cs-cl
24 Jul 2026
← Previous
1…2425262728…161
Next →