AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlog
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
Safety

Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech

DGX agent

arXiv:2608.02941v1 Announce Type: new Abstract: We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-C

safetyarxiv-cs-cl
5 Aug 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Agents

An Actionable Diagnosis of Multilingual, Multi-Agent Planning Failures

DGX agent

arXiv:2608.03735v1 Announce Type: cross Abstract: Multilingual multi-agent systems exhibit substantial degradation beyond English, yet prior work rarely identifies how task-critical information is los

agentsarxiv-cs-cl
5 Aug 2026
Model Releases

ANCHOR-RE: An Agentic Neuro-Symbolic Framework for Grounded Biomedical Relation Extraction

DGX agent

arXiv:2608.03154v1 Announce Type: new Abstract: Biomedical relation extraction (BioRE) extracts structured knowledge from biomedical literature for applications such as knowledge base construction and

model-releasesarxiv-cs-cl
5 Aug 2026
Research

AnchorKV: Anchor-Residual KV Cache Compression

DGX agent

arXiv:2608.02901v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction me

researcharxiv-cs-cl
5 Aug 2026
Model Releases

ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts

DGX agent

arXiv:2608.03898v1 Announce Type: new Abstract: The automatic structural analysis of legal texts is a cornerstone of legal technology, yet the extraction of their logical components remains a signific

model-releasesarxiv-cs-cl
5 Aug 2026
Research

ARCHead: Activation-Metric Residual Correction for Large Language Model Output Heads

DGX agent

arXiv:2608.02703v1 Announce Type: new Abstract: Weight-only quantization substantially reduces the storage of large language model (LLM) transformer blocks, but practical backends often retain the fin

researcharxiv-cs-cl
5 Aug 2026
Model Releases

ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models

DGX agent

arXiv:2608.03358v1 Announce Type: new Abstract: Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emot

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

ATFlash: Per-RoPE-Wavelength Attention Windows for Compute/Memory-Efficient LLM Inference

DGX agent

arXiv:2608.02947v1 Announce Type: cross Abstract: The attention score with rotary position embeddings (RoPE) decomposes exactly into a sum over its 2D-rotation frequency pairs, and each pair's wavelen

model-releasesarxiv-cs-cl
5 Aug 2026
Research

Attention is Case-Sensitive

DGX agent

arXiv:2608.03711v1 Announce Type: cross Abstract: In human visual perception, uppercase lettering serves as a natural salience cue that captures attention within lowercase text. In this paper, we pres

researcharxiv-cs-cl
5 Aug 2026
Model Releases

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

DGX agent

arXiv:2608.03884v1 Announce Type: cross Abstract: In-the-wild Bengali scene text recognition is largely unmeasured: existing resources target handwritten documents or constrained sign-board parsing, r

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

BBOWP-Bench: Evaluating LLMs on Black-Box Optimization Word Problems

DGX agent

arXiv:2608.02612v1 Announce Type: new Abstract: Formulating an optimization problem strongly affects the quality of the final solution, yet good formulations usually require substantial expertise. Rec

model-releasesarxiv-cs-cl
5 Aug 2026
Applications

Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks

DGX agent

arXiv:2608.03340v1 Announce Type: new Abstract: Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream perfor

applicationsarxiv-cs-cl
5 Aug 2026
Research

Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models

DGX agent

arXiv:2608.03038v1 Announce Type: new Abstract: Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how m

researcharxiv-cs-cl
5 Aug 2026
Model Releases

Beyond Initialization Loss: A Systematic Study of Token Embedding Initialization Strategies for LLM Vocabulary Extension

DGX agent

arXiv:2608.03494v1 Announce Type: new Abstract: Vocabulary extension is an efficient way to adapt pretrained large language models (LLMs) to new languages, but the initialization of newly added token

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety

DGX agent

arXiv:2601.17003v2 Announce Type: replace-cross Abstract: Mental-health AI safety is typically evaluated with small, simulation-based benchmarks that may not reflect the linguistic and contextual dive

model-releasesarxiv-cs-cl
5 Aug 2026
Safety

BOW: Training Language Models to Reason Over Plausible Next Words

DGX agent

arXiv:2506.13502v3 Announce Type: replace Abstract: Next-word prediction (NWP) trains language models against a single observed continuation, even though many contexts admit multiple plausible next wo

safetyarxiv-cs-cl
5 Aug 2026
Research

Calibrating Semantic Uncertainty from Observable Language-Model Probabilities

DGX agent

arXiv:2607.17447v2 Announce Type: replace-cross Abstract: As generative artificial intelligence enters scientific and professional work, its uncertainty must be defined on the states that matter for i

researcharxiv-cs-cl
5 Aug 2026
Research

Character Iconicity vs. Arbitrariness: An Arabic NLP Perspective

DGX agent

arXiv:2608.02935v1 Announce Type: new Abstract: Arabic script uses 28 letters, many of which share a common base shape (rasm) and are distinguished only by dot placement. Because early Arabic manuscri

researcharxiv-cs-cl
5 Aug 2026
Local Ai

CIGTSurv: Clinical Information Guided Tri-modal Survival Prediction with Local Prototype Association and Global Feature Alignment

DGX agent

arXiv:2608.03247v1 Announce Type: cross Abstract: Multimodal learning has significantly advanced survival prediction by integrating pathology images with genomic data. However, clinical information, d

local-aiarxiv-cs-cl
5 Aug 2026
Model Releases

ConlangBench: Exploring Language Knowledge and Learning in LLMs through Diverse Constructed Languages

DGX agent

arXiv:2608.03505v1 Announce Type: new Abstract: Constructed languages (conlangs) are intentionally created human languages with a rich tradition of linguistic creativity. Despite their potential for s

model-releasesarxiv-cs-cl
5 Aug 2026
Research

Consensus Measures for Unstructured Biomedical Text Annotations

DGX agent

arXiv:2608.03529v1 Announce Type: new Abstract: Biomedical literature is increasingly mined for knowledge beyond the questions it was written to answer. Because the target concepts are not known in ad

researcharxiv-cs-cl
5 Aug 2026
Research

Convex-Hull-Neighborhood Smooth Dual Generalization: Controlling Local Correction Propagation in Offline RL

DGX agent

arXiv:2608.03108v1 Announce Type: cross Abstract: Offline reinforcement learning (offline RL) can benefit from nearby out-of-distribution (OOD) actions, but estimation errors at these actions may be a

researcharxiv-cs-cl
5 Aug 2026
Applications

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation

DGX agent

arXiv:2608.02694v1 Announce Type: new Abstract: Long-horizon video editing agents receive final-product feedback only after many interdependent decisions. Yet editing quality is subjective, admits mul

applicationsarxiv-cs-cl
5 Aug 2026
Model Releases

Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili

DGX agent

arXiv:2608.03532v1 Announce Type: new Abstract: Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Detecting Hallucinations and Recovering Verified Answers in Arabic Islamic Question Answering

DGX agent

arXiv:2608.03720v1 Announce Type: new Abstract: Large language models can generate fluent responses to Islamic questions while introducing factual errors that are difficult to identify. This paper pre

model-releasesarxiv-cs-cl
5 Aug 2026
Safety

Disentangling Language Modeling and Boundaries

DGX agent

arXiv:2608.03599v1 Announce Type: new Abstract: Byte-level language models are usually argued for on the grounds of robustness, multilingual fairness, and character-level skills. We point to a differe

safetyarxiv-cs-cl
5 Aug 2026
Model Releases

Disentangling MLP Neuron Weights in Vocabulary Space

DGX agent

arXiv:2604.06005v2 Announce Type: replace Abstract: Interpreting the information encoded in language model weights remains a fundamental challenge in mechanistic interpretability. In this work, we int

model-releasesarxiv-cs-cl
5 Aug 2026
Research

Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference

DGX agent

arXiv:2608.03388v1 Announce Type: new Abstract: Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language

researcharxiv-cs-cl
5 Aug 2026
Model Releases

Don't Walk the Line: Boundary Guidance for Filtered Generation

DGX agent

arXiv:2510.11834v3 Announce Type: replace-cross Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tun

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

DP-MemView: A Memory Interface for Attribute-Level Transcript Privacy in Long-Term LLM Agents

DGX agent

arXiv:2608.03130v1 Announce Type: cross Abstract: Long-term memory enables persistent personalization in LLM agents, but repeated memory-conditioned responses can cumulatively reveal protected attribu

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

DS@GT-ARC at eRisk 2026 Task 3: Sparse, Semantic, and LLM Reranking for ADHD Symptom Sentences

DGX agent

arXiv:2608.03883v1 Announce Type: new Abstract: This paper describes our submissions to eRisk 2026 Task 3, ADHD Symptom Sentence Ranking. The task requires systems to rank candidate Reddit sentences a

model-releasesarxiv-cs-cl
5 Aug 2026
Research

DUD: Decoupled Update Dynamics for Reliable Uncertainty Quantification in Large Language Models

DGX agent

arXiv:2608.03411v1 Announce Type: new Abstract: Accurate Uncertainty Quantification (UQ) is critical for reliable deployment of Large Language Models (LLMs), yet traditional probability-based metrics

researcharxiv-cs-cl
5 Aug 2026
Model Releases

Dynamically Allocating Evaluation Effort for Model Ranking

DGX agent

arXiv:2608.03437v1 Announce Type: new Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability. When identifying top-performing m

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Efficient Multilingual Neural Machine Translation via Corpus-Driven Vocabulary Pruning: An English-Arabic Case Study

DGX agent

arXiv:2608.03480v1 Announce Type: new Abstract: The adoption of large pre-trained multilingual models for neural machine translation (MNMT) faces a major challenge: excessive memory and computational

model-releasesarxiv-cs-cl
5 Aug 2026
Research

Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks

DGX agent

arXiv:2608.02966v1 Announce Type: new Abstract: Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scorin

researcharxiv-cs-cl
5 Aug 2026
Model Releases

FLARE: Few-shot Learning-based Adaptive Reflective Engine

DGX agent

arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-

model-releasesarxiv-cs-cl
5 Aug 2026
Tutorials

From SQL Errors to Concept Gaps: An AI-Powered Knowledge Graph Analytics Platform for Personalized Feedback

DGX agent

arXiv:2608.03118v1 Announce Type: new Abstract: This innovative practice full paper describes an AI-powered knowledge graph platform that connects SQL errors to conceptual gaps in undergraduate and gr

tutorialsarxiv-cs-cl
5 Aug 2026
Local Ai

GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models

DGX agent

arXiv:2608.03083v1 Announce Type: cross Abstract: Video large language models (VideoLLMs) achieve strong video understanding performance, but their inference remains expensive due to the large number

local-aiarxiv-cs-cl
5 Aug 2026
Research

HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification

DGX agent

arXiv:2608.03966v1 Announce Type: new Abstract: Large language models can generate fluent Arabic answers while introducing factual errors that are difficult to identify and verify. Existing Arabic hal

researcharxiv-cs-cl
5 Aug 2026
Safety

Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning

DGX agent

arXiv:2608.03545v1 Announce Type: new Abstract: Test-time reinforcement learning (TTRL) improves the reasoning capabilities of large language models without labeled data by updating the policy with ps

safetyarxiv-cs-cl
5 Aug 2026
Model Releases

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

DGX agent

arXiv:2509.23690v2 Announce Type: replace-cross Abstract: Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a ho

model-releasesarxiv-cs-cl
5 Aug 2026
Safety

HomoEnsNER: Does Language Alignment Outperform Architectural Complexity in Gujarati Named Entity Recognition?

DGX agent

arXiv:2608.03105v1 Announce Type: new Abstract: Named Entity Recognition (NER) for Gujarati remains underexplored, hindered by the absence of capitalization cues, rich morphology, lexical ambiguity, a

safetyarxiv-cs-cl
5 Aug 2026
Model Releases

HUKUKBERT: Domain-Specific Language Model for Turkish Law

DGX agent

arXiv:2604.04790v2 Announce Type: replace Abstract: Natural language processing (NLP) advances have powered a generation of LegalTech systems, but Turkish law remains under-served by domain-specific d

model-releasesarxiv-cs-cl
5 Aug 2026
Safety

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

DGX agent

arXiv:2608.03210v1 Announce Type: new Abstract: Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift

safetyarxiv-cs-cl
5 Aug 2026
Model Releases

JudgeArena: A Unified Framework for Reproducible LLM-Judge Evaluation

DGX agent

arXiv:2608.02620v1 Announce Type: new Abstract: LLM-as-a-judge evaluation has become a dominant paradigm for ranking language models, yet the ecosystem remains fragmented: most benchmarks ship their o

model-releasesarxiv-cs-cl
5 Aug 2026
Agents

LACE: Large Language Model Aided Multi-Agent Framework for Agile RISC-V Instruction Extension

DGX agent

arXiv:2608.02915v1 Announce Type: cross Abstract: Domain-specific Instruction Set Architecture eXtensions (ISAX) are widely adopted in the RISC-V ecosystem to accelerate emerging workloads, but implem

agentsarxiv-cs-cl
5 Aug 2026
Research

Language Models Encode the Contextual Truth of Propositions

DGX agent

arXiv:2608.03035v1 Announce Type: new Abstract: Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representati

researcharxiv-cs-cl
5 Aug 2026
Safety

Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR

DGX agent

arXiv:2608.03610v1 Announce Type: new Abstract: Modern LLM-based ASR systems have established multilingual capability as a standard feature, leveraging large-scale multilingual corpora and LLMs' cross

safetyarxiv-cs-cl
5 Aug 2026
← Previous
1…7891011…160
Next →