AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlog
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Research

Evaluating Dedicated Monolingual and Joint Multilingual Causal Models for Dravidian Languages

DGX agent

arXiv:2608.07727v1 Announce Type: new Abstract: Dravidian languages, mainly Tamil, Telugu, Kannada, and Malayalam make up only a small part of the data used to train multilingual language models, so i

researcharxiv-cs-cl
11 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Local Ai

Evidence-Calibrated Runtime Reconstruction for Agent Skills Across Heterogeneous Coding Agents

DGX agent

arXiv:2608.08793v1 Announce Type: new Abstract: Agent Skills package reusable instructions and assets for tool-using language-model agents. Progressive loading creates failure boundaries poorly repres

local-aiarxiv-cs-cl
11 Aug 2026
Model Releases

Evo-Bench: Can Language Models Improve Agent Harness?

DGX agent

arXiv:2608.09096v1 Announce Type: new Abstract: Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emergi

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

EvoTrustRAG: Evolution-Aware Conflict Attribution and Evidence Handling for Reliable Retrieval-Augmented Generation

DGX agent

arXiv:2608.07933v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) improves the factuality of large language models with external knowledge, yet conflicting evidence remains a fund

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Explicit Boundary Markers for Subword Vocabularies

DGX agent

arXiv:2608.08847v1 Announce Type: new Abstract: Subword tokenizers represent many common words twice in space-using writing systems, once with a leading space and once without. The two entries have se

researcharxiv-cs-cl
11 Aug 2026
Research

Failure-Aware Long-Form Translation: Design and Implementation of a Recoverable LLM Translation System

DGX agent

arXiv:2608.09187v1 Announce Type: new Abstract: A long-form translation request can succeed at the API layer and still produce an unusable result. The output may be empty, truncated, filtered, dominat

researcharxiv-cs-cl
11 Aug 2026
Model Releases

FanarGuard: A Culturally-Aware Moderation Filter for Arabic Language Models

DGX agent

arXiv:2511.18852v2 Announce Type: replace Abstract: Content moderation filters are a critical safeguard against alignment failures in language models. Yet most existing filters focus narrowly on gener

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Focus particles and scalar inferences across humans and language models

DGX agent

arXiv:2608.08227v1 Announce Type: new Abstract: Focus particles such as 'even' and 'only' are central to formal semantic theories that posit structured representations over sets of alternatives. 'Even

researcharxiv-cs-cl
11 Aug 2026
Research

From Chains to DAGs: Probing the Graph Structure of Reasoning in LLMs

DGX agent

arXiv:2601.17593v3 Announce Type: replace Abstract: Recent progress in large language models has renewed interest in how multi-step reasoning is represented internally. While prior work often treats r

researcharxiv-cs-cl
11 Aug 2026
Safety

From Rebound to Remedy: Understanding and Mitigating Reward Hacking via Representation Engineering

DGX agent

arXiv:2604.01476v2 Announce Type: replace-cross Abstract: Reinforcement learning for LLMs is vulnerable to reward hacking, where models exploit shortcuts to maximize reward without solving the intende

safetyarxiv-cs-cl
11 Aug 2026
Research

From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios

DGX agent

arXiv:2608.08510v1 Announce Type: new Abstract: Humans have the remarkable ability to engage in spontaneous informal conversations and selectively attend to individual speakers while filtering out com

researcharxiv-cs-cl
11 Aug 2026
Research

From token probabilities to calibrated confidence: An empirical study of mathematical question answering

DGX agent

arXiv:2608.07827v1 Announce Type: cross Abstract: Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns t

researcharxiv-cs-cl
11 Aug 2026
Agents

GraphWalker: Agentic Knowledge Graph Question Answering via Synthetic Trajectory Curriculum

DGX agent

arXiv:2603.28533v3 Announce Type: replace Abstract: Agentic knowledge graph question answering (KGQA) requires an agent to iteratively interact with knowledge graphs (KGs), posing challenges in both t

agentsarxiv-cs-cl
11 Aug 2026
Model Releases

High-Layer Attention Pruning with Rescaling

DGX agent

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional

model-releasesarxiv-cs-cl
11 Aug 2026
Research

HindsightBench: A Black-Box Behavioral Audit Protocol for Parametric Hindsight in Time-Indexed LLM Decision Tasks

DGX agent

arXiv:2607.18867v2 Announce Type: replace-cross Abstract: Large language models leak parametric knowledge of what followed a historical date into decision tasks indexed by that date -- not necessarily

researcharxiv-cs-cl
11 Aug 2026
Research

Human-Level Text-to-SQL via Reinforcement Learning on Verified Data, Without Pipeline Engineering

DGX agent

arXiv:2603.20004v3 Announce Type: replace-cross Abstract: Translating natural language questions to SQL queries (Text-to-SQL) is a long-standing problem in database research. Recent efforts have focus

researcharxiv-cs-cl
11 Aug 2026
Agents

IDRAAK: From Multi-Agent NLP to Few-Shot Prompting for Semantic Drift Detection in Technical Requirements

DGX agent

arXiv:2608.08801v1 Announce Type: new Abstract: Translating technical requirements across languages can introduce semantic drift, altering numerical constraints, polarities, modalities, or other speci

agentsarxiv-cs-cl
11 Aug 2026
Agents

InfMem: Learning System-2 Memory Control for Long-Context Agent

DGX agent

arXiv:2602.02704v2 Announce Type: replace Abstract: Reasoning over ultra-long documents requires synthesizing sparse evidence scattered across distant segments under strict memory constraints. While s

agentsarxiv-cs-cl
11 Aug 2026
Model Releases

Instability of LLM Pre-Pretraining: It Doesn't Always Help. An Investigation on Multiple Languages

DGX agent

arXiv:2608.08800v1 Announce Type: new Abstract: Pretraining LLMs on artificial languages ('pre-pretraining') is a technique that could reportedly increase token efficiency by 33%, i.e., save up to 33%

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation

DGX agent

arXiv:2608.09420v1 Announce Type: new Abstract: User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently

safetyarxiv-cs-cl
11 Aug 2026
Safety

Investigating Multimodal Informativity under Different Partner Visibility Conditions in Video-Mediated Dialogue

DGX agent

arXiv:2608.08915v1 Announce Type: new Abstract: Situated language use is multimodal and embodied. For example, gestures can carry information that is absent or underspecified in the speech signal, yet

safetyarxiv-cs-cl
11 Aug 2026
Safety

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

DGX agent

arXiv:2604.05971v2 Announce Type: replace-cross Abstract: Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

DGX agent

arXiv:2608.09280v1 Announce Type: new Abstract: Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promo

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Jako Tako or Fluent? Presenting PoVisLE: A Polish Vision-Language Evaluation

DGX agent

arXiv:2608.07763v1 Announce Type: new Abstract: Vision-language models (VLMs) have achieved strong performance on tasks such as image captioning, visual question answering, and image-to-text generatio

model-releasesarxiv-cs-cl
11 Aug 2026
Agents

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

DGX agent

arXiv:2608.08255v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge i

agentsarxiv-cs-cl
11 Aug 2026
Model Releases

LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

DGX agent

arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking

DGX agent

arXiv:2608.09650v1 Announce Type: cross Abstract: Reranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial l

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

LogicIF: Towards Complex Logic Instruction Following

DGX agent

arXiv:2508.09125v4 Announce Type: replace Abstract: Instruction following has catalyzed the recent era of Large Language Models (LLMs) and is the foundational skill underpinning more advanced capabili

model-releasesarxiv-cs-cl
11 Aug 2026
Local Ai

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

DGX agent

arXiv:2608.09819v1 Announce Type: cross Abstract: Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after de

local-aiarxiv-cs-cl
11 Aug 2026
Research

'Many Are My Names': The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

DGX agent

arXiv:2608.07852v1 Announce Type: new Abstract: How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexp

researcharxiv-cs-cl
11 Aug 2026
Model Releases

Mawqif-v2: An Arabic Benchmark Dataset for Cross-Target Stance Detection

DGX agent

arXiv:2608.09539v1 Announce Type: new Abstract: Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This pa

model-releasesarxiv-cs-cl
11 Aug 2026
Research

MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL

DGX agent

arXiv:2608.09588v1 Announce Type: new Abstract: Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, het

researcharxiv-cs-cl
11 Aug 2026
Model Releases

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

DGX agent

arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Memorization Dynamics in Knowledge Distillation for Language Models

DGX agent

arXiv:2601.15394v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improv

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy

DGX agent

arXiv:2512.09636v3 Announce Type: replace Abstract: Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, an

model-releasesarxiv-cs-cl
11 Aug 2026
Research

MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs

DGX agent

arXiv:2510.19366v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) scales model capacity through sparse activation, and is becoming an important architecture for large language models (LLMs)

researcharxiv-cs-cl
11 Aug 2026
Research

Multilingual Emotion Neurons in Large Audio-Language Models

DGX agent

arXiv:2608.08772v1 Announce Type: new Abstract: Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on

researcharxiv-cs-cl
11 Aug 2026
Safety

North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings

DGX agent

arXiv:2608.08607v1 Announce Type: new Abstract: Mental health disorders are a leading cause of disability worldwide, yet Natural Language Processing (NLP) research for mental healthcare has remained c

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

DGX agent

arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost ex

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

On the use of foundation models in cognitive science

DGX agent

arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations o

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

DGX agent

arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

PAM: Training Policy-Aligned Moderation Filters at Scale

DGX agent

arXiv:2505.19766v4 Announce Type: replace Abstract: Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet exi

safetyarxiv-cs-cl
11 Aug 2026
Safety

Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics

DGX agent

arXiv:2510.22028v4 Announce Type: replace Abstract: Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data

safetyarxiv-cs-cl
11 Aug 2026
Safety

Position Bias in Ordinal Classification: A Systematic Evaluation

DGX agent

arXiv:2608.08869v1 Announce Type: new Abstract: Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predi

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

DGX agent

arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

DGX agent

arXiv:2608.09551v1 Announce Type: new Abstract: In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural lang

safetyarxiv-cs-cl
11 Aug 2026
Research

PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings

DGX agent

arXiv:2608.07478v1 Announce Type: cross Abstract: India's 22 official languages create a critical accessibility barrier: the majority of medical documentation exists exclusively in English, yet the pa

researcharxiv-cs-cl
11 Aug 2026
Research

PRISM-Delta: Differential Subspace Steering for Prompt Highlighting in Large Language Models

DGX agent

arXiv:2603.10705v2 Announce Type: replace Abstract: Prompt highlighting steers a large language model to prioritize user-specified text spans during generation. A key challenge of existing Key-editing

researcharxiv-cs-cl
11 Aug 2026
← Previous
12345…160
Next →