AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
11 Aug 2026

Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning

AgentsDGX agent

arXiv:2608.08255v1 Announce Type: cross Abstract: Agentic reinforcement learning (RL) often suffers from delayed and sparse rewards in real-world environments. A promising solution to this challenge i

LexKairos: Benchmarking Legal Temporal Capabilities in LLMs

Model ReleasesDGX agent

arXiv:2608.09106v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated strong performance across a wide range of legal tasks. In legal practice, time is a critical concept that

Listwise Cross-Encoder Fine-Tuning vs. Agentic Instruction Tuning for LLM Rerankers: A Systematic Study in Medical Procedure Reranking

Model Releases

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
DGX agent

arXiv:2608.09650v1 Announce Type: cross Abstract: Reranking medical procedures against patient queries is a critical component of health insurance information retrieval, complicated by a substantial l

LogicIF: Towards Complex Logic Instruction Following

Model ReleasesDGX agent

arXiv:2508.09125v4 Announce Type: replace Abstract: Instruction following has catalyzed the recent era of Large Language Models (LLMs) and is the foundational skill underpinning more advanced capabili

Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Local AiDGX agent

arXiv:2608.09819v1 Announce Type: cross Abstract: Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after de

'Many Are My Names': The Anatomy of the Assistant and Its Personas via Sparse Autoencoders

ResearchDGX agent

arXiv:2608.07852v1 Announce Type: new Abstract: How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexp

Mawqif-v2: An Arabic Benchmark Dataset for Cross-Target Stance Detection

Model ReleasesDGX agent

arXiv:2608.09539v1 Announce Type: new Abstract: Publicly available Arabic datasets for target-specific stance detection remain limited, particularly for evaluating cross-target generalization. This pa

MDB-Link: Hierarchical Schema Linking for Multi-Database Text-to-SQL

ResearchDGX agent

arXiv:2608.09588v1 Announce Type: new Abstract: Traditional Text-to-SQL research and benchmarks assume a known target database, overlooking settings in which a query must be routed within a large, het

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

Model ReleasesDGX agent

arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe

Memorization Dynamics in Knowledge Distillation for Language Models

Model ReleasesDGX agent

arXiv:2601.15394v2 Announce Type: replace Abstract: Knowledge Distillation (KD) is increasingly adopted to transfer capabilities from large language models to smaller ones, offering significant improv

MiraMind: Benchmarking Reliable Mental Health Reasoning beyond Answer Accuracy

Model ReleasesDGX agent

arXiv:2512.09636v3 Announce Type: replace Abstract: Mental-health reasoning with large language models (LLMs) is an evidence-constrained judgment problem: models must transform limited, subjective, an

MoE-Prism: Disentangling Monolithic Experts for Elastic MoE Services via Model-System Co-Designs

ResearchDGX agent

arXiv:2510.19366v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) scales model capacity through sparse activation, and is becoming an important architecture for large language models (LLMs)

Multilingual Emotion Neurons in Large Audio-Language Models

ResearchDGX agent

arXiv:2608.08772v1 Announce Type: new Abstract: Emotion is central to human communication, and its expression varies across languages. Large audio-language models (LALMs) achieve strong performance on

North Africa's Missing Framework: NLP-Driven Mental Healthcare in Algeria and Implications for Low-resource Settings

SafetyDGX agent

arXiv:2608.08607v1 Announce Type: new Abstract: Mental health disorders are a leading cause of disability worldwide, yet Natural Language Processing (NLP) research for mental healthcare has remained c

OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents

Model ReleasesDGX agent

arXiv:2608.08775v1 Announce Type: new Abstract: Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost ex

On the use of foundation models in cognitive science

SafetyDGX agent

arXiv:2608.07812v1 Announce Type: new Abstract: A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations o

OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories

Model ReleasesDGX agent

arXiv:2608.08557v1 Announce Type: new Abstract: Visual tool use has emerged as a fundamental capability for multimodal agents to actively acquire evidence beyond a fixed image encoding. The prevailing

PAM: Training Policy-Aligned Moderation Filters at Scale

SafetyDGX agent

arXiv:2505.19766v4 Announce Type: replace Abstract: Large language models (LLMs) remain vulnerable to misalignment and jailbreaks, making external safeguards like moderation filters essential, yet exi

Penalizing Length: Uncovering Systematic Bias in Quality Estimation Metrics

SafetyDGX agent

arXiv:2510.22028v4 Announce Type: replace Abstract: Quality Estimation (QE) metrics are vital in machine translation for reference-free evaluation and increasingly serve as selection criteria in data

Position Bias in Ordinal Classification: A Systematic Evaluation

SafetyDGX agent

arXiv:2608.08869v1 Announce Type: new Abstract: Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predi

PragMatch: Separating Pragmatic Incongruity from Cross-Modal Mismatch in Large Vision-Language Models

Model ReleasesDGX agent

arXiv:2608.09772v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) have demonstrated strong performance on multimodal benchmarks, yet it remains unclear whether they genuinely reason

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

SafetyDGX agent

arXiv:2608.09551v1 Announce Type: new Abstract: In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural lang

PragyaDoc: A Universal Document Intelligence Framework for Multilingual Medical Document Understanding in Low-Resource Settings

ResearchDGX agent

arXiv:2608.07478v1 Announce Type: cross Abstract: India's 22 official languages create a critical accessibility barrier: the majority of medical documentation exists exclusively in English, yet the pa

PRISM-Delta: Differential Subspace Steering for Prompt Highlighting in Large Language Models

ResearchDGX agent

arXiv:2603.10705v2 Announce Type: replace Abstract: Prompt highlighting steers a large language model to prioritize user-specified text spans during generation. A key challenge of existing Key-editing

Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

Model ReleasesDGX agent

arXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision

R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning

ResearchDGX agent

arXiv:2602.05940v2 Announce Type: replace Abstract: Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reas

RA-FinBERT: Rule-aware LoRA adaptation for low-resource financial sentiment classification

Model ReleasesDGX agent

arXiv:2608.09834v1 Announce Type: new Abstract: Financial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making. Existi

Reading Cognition as Decisions Unfold in Words: A Factorized Inverse Decision Model

ApplicationsDGX agent

arXiv:2608.09222v1 Announce Type: new Abstract: Inverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action traj

Reducing Pretraining-Generation Mismatch in Diffusion Language Models

Model ReleasesDGX agent

arXiv:2608.09424v1 Announce Type: new Abstract: Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left cont

REFRAMED: Towards Realistic Audio Description Generation for Movies

Model ReleasesDGX agent

arXiv:2608.09765v1 Announce Type: new Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video cap

Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models

SafetyDGX agent

arXiv:2508.16406v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature

Routesplain: Towards Faithful and Intervenable Routing for Software-related Tasks

ResearchDGX agent

arXiv:2511.09373v2 Announce Type: replace-cross Abstract: LLMs now tackle a wide range of software-related tasks, yet we show that their performance varies markedly both across and within these tasks.

Safety Cost of Steering Vectors Is Separable and Reducible

SafetyDGX agent

arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally comprom

SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems

Model ReleasesDGX agent

arXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos

Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access

Model ReleasesDGX agent

arXiv:2608.08942v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide

SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex Computation

Model ReleasesDGX agent

arXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To en

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

ResearchDGX agent

arXiv:2608.09568v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing

Security and Privacy Taxonomy Generation from Mobile App Reviews

ResearchDGX agent

arXiv:2608.09049v1 Announce Type: new Abstract: Mobile app reviews are a rich, continuously renewing source of how users experience privacy and security, yet existing taxonomies of these concerns are

Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents

Model ReleasesDGX agent

arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this c

StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning

HardwareDGX agent

arXiv:2603.02637v2 Announce Type: replace-cross Abstract: Modern machine learning (ML) workloads increasingly rely on GPUs, yet achieving high end-to-end performance remains challenging due to depende

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

ResearchDGX agent

arXiv:2608.09767v1 Announce Type: new Abstract: Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological

Subjective Multi-Bias Detection with Large Language Models

Model ReleasesDGX agent

arXiv:2608.09126v1 Announce Type: new Abstract: In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identificati

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

Model ReleasesDGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

Model ReleasesDGX agent

arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

Model ReleasesDGX agent

arXiv:2608.09802v1 Announce Type: new Abstract: As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluati

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

Model ReleasesDGX agent

arXiv:2608.07851v1 Announce Type: cross Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expre

Testing Hypotheses from the Social Approval Theory of Online Hate: An Analysis of 110 Million Messages from Parler

ResearchDGX agent

arXiv:2507.10810v3 Announce Type: replace Abstract: We examined how social approval motivates online hate via the social approval theory, which argues social approval signals on hate messages predict

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

Model ReleasesDGX agent

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

Model ReleasesDGX agent

arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution c

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

Model ReleasesDGX agent

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

The No-Meaning Falsity: The Structural Impossibility of the Arbitrary Sign in Classical Arabic

ResearchDGX agent

arXiv:2608.07737v1 Announce Type: new Abstract: This paper investigates whether the postmodern claim of unrestricted semantic indeterminacy, and its foundational Saussurean axiom of the arbitrary sign

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

AgentsDGX agent

arXiv:2608.08239v1 Announce Type: cross Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents

The Transparency Trap: How AI Disclaimers Create Overconfidence in High-Stakes Decisions

SafetyDGX agent

arXiv:2608.07493v1 Announce Type: cross Abstract: Current AI disclaimers often fail to function as intended due to warning habituation and a transparency paradox. As AI-generated information becomes p

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

SafetyDGX agent

arXiv:2608.07980v1 Announce Type: cross Abstract: In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the

Thinking Is Not Telling: Information Disclosure in User-Service LLM Agents

AgentsDGX agent

arXiv:2602.07796v2 Announce Type: replace Abstract: User-engaged LLM agents increasingly operate in service scenarios where task success depends on coordination between the agent, the user, and a stat

Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

Model ReleasesDGX agent

arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this e

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

Model ReleasesDGX agent

arXiv:2608.08910v1 Announce Type: new Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses

Tokenisation over Bounded Alphabets is Hard

ResearchDGX agent

arXiv:2511.15709v2 Announce Type: replace Abstract: Recent works have shown that tokenisation is NP-complete. However, these works assume tokenisation is applied to inputs with unboundedly large alpha

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

Model ReleasesDGX agent

arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o

Towards an LLM-based method for quantifying the sexual content in song lyrics

Model ReleasesDGX agent

arXiv:2608.08885v1 Announce Type: cross Abstract: Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mo

← Previous
12345…128
Next →