AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Model Releases

Prompt engineering does not universally improve Large Language Model performance across clinical decision-making tasks

DGX agent

arXiv:2512.22966v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated promise in medical knowledge assessments, yet their practical utility in real-world clinical decision

model-releasesarxiv-cs-cl
11 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Research

R3S: Refining and Recovering Reinforcement Signals for Multilingual Understanding and Reasoning

DGX agent

arXiv:2602.05940v2 Announce Type: replace Abstract: Large reasoning models often default to English reasoning when processing non-English questions, yet their performance drops substantially when reas

researcharxiv-cs-cl
11 Aug 2026
Model Releases

RA-FinBERT: Rule-aware LoRA adaptation for low-resource financial sentiment classification

DGX agent

arXiv:2608.09834v1 Announce Type: new Abstract: Financial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making. Existi

model-releasesarxiv-cs-cl
11 Aug 2026
Applications

Reading Cognition as Decisions Unfold in Words: A Factorized Inverse Decision Model

DGX agent

arXiv:2608.09222v1 Announce Type: new Abstract: Inverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action traj

applicationsarxiv-cs-cl
11 Aug 2026
Model Releases

Reducing Pretraining-Generation Mismatch in Diffusion Language Models

DGX agent

arXiv:2608.09424v1 Announce Type: new Abstract: Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left cont

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

REFRAMED: Towards Realistic Audio Description Generation for Movies

DGX agent

arXiv:2608.09765v1 Announce Type: new Abstract: Audio Description (AD) is a verbal narration of key visual content in videos, enabling access for visually impaired audiences. Unlike standard video cap

model-releasesarxiv-cs-cl
11 Aug 2026
Safety

Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models

DGX agent

arXiv:2508.16406v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs. The evolving nature

safetyarxiv-cs-cl
11 Aug 2026
Research

Routesplain: Towards Faithful and Intervenable Routing for Software-related Tasks

DGX agent

arXiv:2511.09373v2 Announce Type: replace-cross Abstract: LLMs now tackle a wide range of software-related tasks, yet we show that their performance varies markedly both across and within these tasks.

researcharxiv-cs-cl
11 Aug 2026
Safety

Safety Cost of Steering Vectors Is Separable and Reducible

DGX agent

arXiv:2608.08383v1 Announce Type: new Abstract: Steering vectors are a lightweight tool for controlling LLM behavior. However, emerging evidence shows that steering vectors can unintentionally comprom

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems

DGX agent

arXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cos

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access

DGX agent

arXiv:2608.08942v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex Computation

DGX agent

arXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To en

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Se-DPO: Self-Evolving Token Credit for Direct Preference Optimization

DGX agent

arXiv:2608.09568v1 Announce Type: new Abstract: Direct Preference Optimization (DPO) aggregates token-level log-probability ratios via uniform summation, implicitly treating all tokens as contributing

researcharxiv-cs-cl
11 Aug 2026
Research

Security and Privacy Taxonomy Generation from Mobile App Reviews

DGX agent

arXiv:2608.09049v1 Announce Type: new Abstract: Mobile app reviews are a rich, continuously renewing source of how users experience privacy and security, yet existing taxonomies of these concerns are

researcharxiv-cs-cl
11 Aug 2026
Model Releases

Sparks of Cooperative Reasoning: LLMs as Strategic Hanabi Agents

DGX agent

arXiv:2601.18077v3 Announce Type: replace Abstract: Cooperative reasoning under incomplete information remains challenging for both humans and multi-agent systems. The card game Hanabi embodies this c

model-releasesarxiv-cs-cl
11 Aug 2026
Hardware

StitchCUDA: An Automated Multi-Agents End-to-End GPU Programing Framework with Rubric-based Agentic Reinforcement Learning

DGX agent

arXiv:2603.02637v2 Announce Type: replace-cross Abstract: Modern machine learning (ML) workloads increasingly rely on GPUs, yet achieving high end-to-end performance remains challenging due to depende

hardwarearxiv-cs-cl
11 Aug 2026
Research

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

DGX agent

arXiv:2608.09767v1 Announce Type: new Abstract: Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological

researcharxiv-cs-cl
11 Aug 2026
Model Releases

Subjective Multi-Bias Detection with Large Language Models

DGX agent

arXiv:2608.09126v1 Announce Type: new Abstract: In this project, we delved into the pervasive challenge of bias detection within the text content. More specifically, our focus lies on the identificati

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

DGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

DGX agent

arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

DGX agent

arXiv:2608.09802v1 Announce Type: new Abstract: As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluati

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

TEMPER: Tensorized Efficient Manifold-constrained Parameterization for Expressive Residual Routing

DGX agent

arXiv:2608.07851v1 Announce Type: cross Abstract: Residual connections rely on a static residual pathway, and are essential for training deep neural networks. Hyper-connections (HC) increase the expre

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Testing Hypotheses from the Social Approval Theory of Online Hate: An Analysis of 110 Million Messages from Parler

DGX agent

arXiv:2507.10810v3 Announce Type: replace Abstract: We examined how social approval motivates online hate via the social approval theory, which argues social approval signals on hate messages predict

researcharxiv-cs-cl
11 Aug 2026
Model Releases

Tevatron-Elastic: A Unified Abstraction for Training Elastic Retrievers and Rerankers

DGX agent

arXiv:2608.08809v1 Announce Type: new Abstract: A single model scale challenges the flexibility of a production retrieval system: some settings need it faster, others need a smaller index, and the rig

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

DGX agent

arXiv:2608.08650v1 Announce Type: new Abstract: Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution c

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction

DGX agent

arXiv:2608.07517v1 Announce Type: cross Abstract: Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aw

model-releasesarxiv-cs-cl
11 Aug 2026
Research

The No-Meaning Falsity: The Structural Impossibility of the Arbitrary Sign in Classical Arabic

DGX agent

arXiv:2608.07737v1 Announce Type: new Abstract: This paper investigates whether the postmodern claim of unrestricted semantic indeterminacy, and its foundational Saussurean axiom of the arbitrary sign

researcharxiv-cs-cl
11 Aug 2026
Agents

The Replay Gap: Static Evaluation of Model Switching in LLM Agents Scores the Wrong World

DGX agent

arXiv:2608.08239v1 Announce Type: cross Abstract: LLM routers promise efficiency by matching each request to the cheapest adequate model, and are increasingly applied per step inside multi-step agents

agentsarxiv-cs-cl
11 Aug 2026
Safety

The Transparency Trap: How AI Disclaimers Create Overconfidence in High-Stakes Decisions

DGX agent

arXiv:2608.07493v1 Announce Type: cross Abstract: Current AI disclaimers often fail to function as intended due to warning habituation and a transparency paradox. As AI-generated information becomes p

safetyarxiv-cs-cl
11 Aug 2026
Safety

The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints

DGX agent

arXiv:2608.07980v1 Announce Type: cross Abstract: In recent years, the term voiceprint has regained attention, particularly in technological applications and policy-making contexts, often carrying the

safetyarxiv-cs-cl
11 Aug 2026
Agents

Thinking Is Not Telling: Information Disclosure in User-Service LLM Agents

DGX agent

arXiv:2602.07796v2 Announce Type: replace Abstract: User-engaged LLM agents increasingly operate in service scenarios where task success depends on coordination between the agent, the user, and a stat

agentsarxiv-cs-cl
11 Aug 2026
Model Releases

Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

DGX agent

arXiv:2608.08168v1 Announce Type: new Abstract: While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this e

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Tied Trit-Planes: Constraining PTQTP to a Uniform Nine-Level Quantizer, with a Persistent Folded Format for Disk-Streamed Mixture-of-Experts Serving

DGX agent

arXiv:2608.08910v1 Announce Type: new Abstract: PTQTP decomposes LLM weight matrices into two ternary (trit) planes with two free per-group scales. Tying the scales to a fixed ratio of three collapses

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Tokenisation over Bounded Alphabets is Hard

DGX agent

arXiv:2511.15709v2 Announce Type: replace Abstract: Recent works have shown that tokenisation is NP-complete. However, these works assume tokenisation is applied to inputs with unboundedly large alpha

researcharxiv-cs-cl
11 Aug 2026
Model Releases

Toward Metacognitive One-Shot Indirect Prompt Injection: Strategy Abstraction Via Outcome-Conditioned Reflection

DGX agent

arXiv:2608.08795v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents are vulnerable to indirect prompt injection (IPI), in which malicious instructions embedded in external o

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Towards an LLM-based method for quantifying the sexual content in song lyrics

DGX agent

arXiv:2608.08885v1 Announce Type: cross Abstract: Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mo

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents

DGX agent

arXiv:2608.09044v1 Announce Type: new Abstract: Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically r

researcharxiv-cs-cl
11 Aug 2026
Safety

Uncertainty-Aware Variational Reward Factorization via Probabilistic Preference Bases for LLM Personalization

DGX agent

arXiv:2604.00997v2 Announce Type: replace Abstract: Reward factorization personalizes large language models (LLMs) by decomposing rewards into shared basis functions and user-specific weights. Yet, ex

safetyarxiv-cs-cl
11 Aug 2026
Applications

Universal or Language-Family-Specific Script Unification for Cross-Lingual Transfer? A Case Study on Turkic Languages

DGX agent

arXiv:2608.09356v1 Announce Type: new Abstract: Closely related languages written in different scripts expose little surface overlap to multilingual models, limiting cross-lingual transfer. We compare

applicationsarxiv-cs-cl
11 Aug 2026
Model Releases

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

DGX agent

arXiv:2608.09209v1 Announce Type: new Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

UNSPECIFIC: General Constraint Synthesis for Breaking Copy-and-Paste Shortcut in LLM Instruction Following

DGX agent

arXiv:2608.09154v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly expected to follow long lists of constraints in complex instructions, and synthesizing instructions from a

model-releasesarxiv-cs-cl
11 Aug 2026
Research

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

DGX agent

arXiv:2608.08791v1 Announce Type: new Abstract: Diffusion language models use broad context to create text, suggesting they might handle input noise better than standard models. Testing reveals this i

researcharxiv-cs-cl
11 Aug 2026
Model Releases

VectraYX-Vision-1B: A Sub-2B Spanish/LATAM Cybersecurity Vision-Language Model with Structured Visual Reasoning and Native Tool Use

DGX agent

arXiv:2608.08477v1 Announce Type: new Abstract: We present VectraYX-Vision-1B, a sub-2B vision-language model (VLM) for Spanish/LATAM cybersecurity imagery, coupling a frozen SigLIP-so400m encoder to

model-releasesarxiv-cs-cl
11 Aug 2026
Local Ai

Verifiably grounded machine interpretation of lunar geology

DGX agent

arXiv:2608.09276v1 Announce Type: new Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations. Here, we present a step toward an a

local-aiarxiv-cs-cl
11 Aug 2026
Research

VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding

DGX agent

arXiv:2608.09698v1 Announce Type: cross Abstract: Great fiction earns its verisimilitude through precise details, from how a longsword is gripped to pierce armor gaps to why a bleeding corpse cannot y

researcharxiv-cs-cl
11 Aug 2026
Research

Wisdom in Unity: The Role of Multilingual Training in Figurative Language Identification in Proverbs

DGX agent

arXiv:2608.08090v1 Announce Type: new Abstract: Although multilingual approaches to figurative language identification are not new, the shift beyond language homogeneous training data requires a clear

researcharxiv-cs-cl
11 Aug 2026
Safety

A Practical Evaluation Method for Long-Form Simultaneous Speech-to-Speech Translation

DGX agent

arXiv:2606.15059v2 Announce Type: replace Abstract: Simultaneous speech-to-speech translation (SimulS2ST) enables real-time cross-lingual communication, but existing evaluation has focused largely on

safetyarxiv-cs-cl
10 Aug 2026
Research

AfriNLLB: Efficient Translation Models for African Languages

DGX agent

arXiv:2602.09373v2 Announce Type: replace Abstract: In this work, we present AfriNLLB, a series of lightweight models for efficient translation from and into African languages. AfriNLLB supports 15 la

researcharxiv-cs-cl
10 Aug 2026
← Previous
123456…160
Next →