AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,700 results
Model Releases

AGC-Bench: Measuring Artificial General Creativity

DGX agent

arXiv:2607.01152v1 Announce Type: new Abstract: Creativity research has debated whether creativity is domain-specific (e.g., visual, writing, science), and if it is psychometrically separable from gen

model-releasesarxiv-cs-cl
2 Jul 2026
Research
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs

DGX agent

arXiv:2607.00171v1 Announce Type: new Abstract: Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a

researcharxiv-cs-cl
2 Jul 2026
Model Releases

Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth

DGX agent

arXiv:2607.00139v1 Announce Type: new Abstract: The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acu

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

Beyond Document Grounding: Span-Level Hallucination Detection over Code, Tool Output, and Documents

DGX agent

arXiv:2607.00895v1 Announce Type: new Abstract: Hallucination detection for retrieval-augmented generation (RAG) is usually evaluated on natural-language document evidence. However, grounded generatio

model-releasesarxiv-cs-cl
2 Jul 2026
Local Ai

Beyond Perplexity: A Behavioral Evaluation Framework for Deployment-Memory Claims in LLM Test-Time Training

DGX agent

arXiv:2607.00368v1 Announce Type: new Abstract: Large language model test-time training (TTT) is often evaluated through local proxy metrics: models are updated on recent tokens, retrieved context, ta

local-aiarxiv-cs-cl
2 Jul 2026
Model Releases

Clinician-Level Agreement Without Clinical Caution: LLM Evaluator Limits in Medical AI Benchmarking

DGX agent

arXiv:2607.01103v1 Announce Type: new Abstract: Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated L

model-releasesarxiv-cs-cl
2 Jul 2026
Tutorials

CogTax: A Four-Level Cognitive Taxonomy for Command-Line Computing Education

DGX agent

arXiv:2607.00140v1 Announce Type: cross Abstract: As computing education expands beyond traditional programming into operational domains such as systems administration and command-line environments, e

tutorialsarxiv-cs-cl
2 Jul 2026
Agents

Conversable Complexity: Agentic LLM Collectives as Interpretable Substrates

DGX agent

arXiv:2607.01047v1 Announce Type: new Abstract: Complexity and interpretability rarely coincide: systems rich enough for complex behaviours to emerge are usually too opaque to question, while transpar

agentsarxiv-cs-cl
2 Jul 2026
Research

Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages

DGX agent

arXiv:2607.01161v1 Announce Type: cross Abstract: Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different

researcharxiv-cs-cl
2 Jul 2026
Research

'Don't Say It!': Constraints, Compliance, and Communication when Language Models Play Taboo

DGX agent

arXiv:2607.00601v1 Announce Type: new Abstract: The game of Taboo requires describing a target word without using a set of forbidden words, so that other players can guess it. This deceptively simple

researcharxiv-cs-cl
2 Jul 2026
Model Releases

Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation

DGX agent

arXiv:2607.00570v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) increasingly requires models to answer questions from multiple retrieved documents, where only some sources are rel

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

Dynamic Bidirectional Pattern Memory: A Production-Scale Empirical Characterisation of Inference-Time Gating in Clinical NLP

DGX agent

arXiv:2607.00870v1 Announce Type: new Abstract: We study inference-time pattern-memory gating in a production-scale clinical natural language processing (NLP) pipeline. The pipeline pairs a generator

model-releasesarxiv-cs-cl
2 Jul 2026
Research

Efficient Multilingual Reasoning Transfer via Progressive Code-Switching

DGX agent

arXiv:2607.00485v1 Announce Type: new Abstract: Large reasoning models (LRMs) have achieved strong reasoning capabilities in English, yet their performance degrades significantly when required to reas

researcharxiv-cs-cl
2 Jul 2026
Model Releases

EPC: A Standardized Protocol for Measuring Evaluator Preference Dynamics in LLM Agent Systems

DGX agent

arXiv:2607.00297v1 Announce Type: cross Abstract: When LLM agents use evaluator feedback to adapt their behavior in closed loops, evaluator biases propagate through the agent's strategy distribution -

model-releasesarxiv-cs-cl
2 Jul 2026
Applications

Evidence-Supported Credit Risk Report Generation Using News-Centric Financial Knowledge Graphs

DGX agent

arXiv:2607.01023v1 Announce Type: new Abstract: Financial markets evolve in response to real-world events reported in news, yet these drivers often remain implicit in text. To better explain market dy

applicationsarxiv-cs-cl
2 Jul 2026
Model Releases

ext{Log}_ext{b}Quant: Quantizing Language Models in Logarithmic Space

DGX agent

arXiv:2607.01127v1 Announce Type: new Abstract: Quantization has become an invaluable tool to reduce memory requirements and inference speed of modern language models, in particular to make them avail

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

From Holistic Evaluation to Structured Criteria: Rubrics Across the Evolving LLM Landscape

DGX agent

arXiv:2606.08625v2 Announce Type: replace Abstract: As Large Language Models (LLMs) advance toward open-ended autonomous agents, the mechanisms used to evaluate and guide their behavior must evolve ac

safetyarxiv-cs-cl
2 Jul 2026
Agents

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

DGX agent

arXiv:2601.04424v2 Announce Type: replace Abstract: Large language models (LLMs) now support contexts of up to 1M tokens, but their strengths and weaknesses on complex long-context tasks remain unclea

agentsarxiv-cs-cl
2 Jul 2026
Model Releases

GPTKB v1.5: A Massive Knowledge Base for Exploring Factual LLM Knowledge

DGX agent

arXiv:2507.05740v2 Announce Type: replace Abstract: Language models are powerful artifacts, yet their factual knowledge is still poorly understood, and inaccessible to ad-hoc browsing and scalable sta

model-releasesarxiv-cs-cl
2 Jul 2026
Research

Graded strength of comparative illusions is explained by Bayesian inference

DGX agent

arXiv:2511.14642v2 Announce Type: replace Abstract: Like visual processing, language processing is susceptible to illusions in which people systematically misperceive stimuli. In one such case--the co

researcharxiv-cs-cl
2 Jul 2026
Research

How Do We Engage with Other Disciplines? A Framework to Study Meaningful Interdisciplinary Discourse in Scholarly Publications

DGX agent

arXiv:2601.17020v2 Announce Type: replace-cross Abstract: With the rising popularity of interdisciplinary work and increasing institutional incentives in this direction, there is a growing need to und

researcharxiv-cs-cl
2 Jul 2026
Research

How Ethos and Pathos Appeals Resonate in Reader Interpretations of Social Media Messages

DGX agent

arXiv:2607.00873v1 Announce Type: new Abstract: Rhetorical strategies and their influence on audiences are often studied through social media posts and comments. However, this focus overlooks the univ

researcharxiv-cs-cl
2 Jul 2026
Model Releases

Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting

DGX agent

arXiv:2607.00159v1 Announce Type: new Abstract: Knowledge-Based Visual Question Answering (KB-VQA) aims to evaluate whether Visual Language Models (VLMs) can retrieve, ground, and reason over external

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

DGX agent

arXiv:2607.01232v1 Announce Type: cross Abstract: Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adapta

model-releasesarxiv-cs-cl
2 Jul 2026
Research

Know When to Stop: Segment-Level Credit Assignment for Reducing Overthinking

DGX agent

arXiv:2607.00482v1 Announce Type: new Abstract: Reasoning language models frequently overthink: generating extended chains of behaviors such as hedging, approach abandonment, and self contradiction th

researcharxiv-cs-cl
2 Jul 2026
Research

KnowledgeDebugger -- an Exploration Tool for Knowledge Localization and Editing in Transformers

DGX agent

arXiv:2607.01000v1 Announce Type: new Abstract: Recent research has increasingly focused on understanding how Transformers store and process knowledge, as well as how this knowledge can be edited. Res

researcharxiv-cs-cl
2 Jul 2026
Research

Low Perplexity is Repetition: A One-Dimensional Self-Conditioning Attractor in Continuous Diffusion LMs

DGX agent

arXiv:2607.00588v1 Announce Type: new Abstract: Continuous diffusion language models such as ELF report record-low generative perplexity (Gen-PPL). We find a catch: these models repeat far more than h

researcharxiv-cs-cl
2 Jul 2026
Model Releases

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

DGX agent

arXiv:2510.24434v3 Announce Type: replace Abstract: The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resource linguistic settings due to a lack of high-quali

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

LV-ROVER: Multi-Stream Tesseract Voting for Maltese Paragraph OCR

DGX agent

arXiv:2607.00250v1 Announce Type: new Abstract: Maltese has decent text corpora and pretrained language models, but, like many languages outside the handful with large OCR benchmarks, only a single kn

model-releasesarxiv-cs-cl
2 Jul 2026
Research

Message Passing Enables Efficient Reasoning

DGX agent

arXiv:2607.01077v1 Announce Type: new Abstract: While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is

researcharxiv-cs-cl
2 Jul 2026
Research

MetaHOPE: A Metaphor-Oriented Evaluation Framework for Analysing MT and LLM Translation Errors

DGX agent

arXiv:2607.00848v1 Announce Type: new Abstract: In this opinion paper, we propose MetaHOPE, an error severity-aware annotation framework for evaluating metaphor translations. Metaphors present challen

researcharxiv-cs-cl
2 Jul 2026
Model Releases

MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules

DGX agent

arXiv:2607.00464v1 Announce Type: cross Abstract: Current molecular generation benchmarks emphasize task complexity, molecule novelty, and property alignment; they largely overlook a critical concern:

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

MSQA: A Natively Sourced Multilingual and Multicultural SimpleQA Benchmark

DGX agent

arXiv:2607.00724v1 Announce Type: new Abstract: Multilingual fluency often invites a stronger assumption: a model that can speak a user's language must also understand the culture encoded by that lang

model-releasesarxiv-cs-cl
2 Jul 2026
Agents

Multi-Turn Agentic Scientific Literature Search via Workflow Induction

DGX agent

arXiv:2607.00597v1 Announce Type: new Abstract: Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, an

agentsarxiv-cs-cl
2 Jul 2026
Model Releases

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

DGX agent

arXiv:2607.00890v1 Announce Type: new Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

OpenReward: Learning to Reward Long-form Agentic Tasks via Reinforcement Learning

DGX agent

arXiv:2510.24636v3 Announce Type: replace Abstract: Reward models (RMs) have become essential for aligning large language models (LLMs), serving as scalable proxies for human evaluation in both traini

safetyarxiv-cs-cl
2 Jul 2026
Research

Persona Non Grata: LLM Persona-Driven Generations in MCQA are Unstable in Distinct Dimensions

DGX agent

arXiv:2607.00937v1 Announce Type: new Abstract: Persona-driven generations (PDGs) have seen prolific use in research and industry applications, where a large language model (LLM) takes on a 'persona'

researcharxiv-cs-cl
2 Jul 2026
Model Releases

Quantifying the Affective Gap: A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion Taxonomies

DGX agent

arXiv:2607.00968v1 Announce Type: new Abstract: Emotion recognition in natural language is a foundational challenge in affective computing, with critical implications for human-computer interaction, m

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling

DGX agent

arXiv:2607.01179v1 Announce Type: cross Abstract: Scaling inference compute, by generating many parallel attempts per problem, is a costly but reliable lever for improving language model capabilities.

safetyarxiv-cs-cl
2 Jul 2026
Local Ai

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination

DGX agent

arXiv:2607.00158v1 Announce Type: new Abstract: Hallucination remains one of the central obstacles to deploying medical LLMs. Yet, even when hallucination can be detected, it is still unclear whether

local-aiarxiv-cs-cl
2 Jul 2026
Research

Rosetta: Composable Native Multimodal Pretraining

DGX agent

arXiv:2607.00293v1 Announce Type: cross Abstract: Achieving true artificial general intelligence requires foundation models capable of integrating new modalities without forgetting prior knowledge. Ho

researcharxiv-cs-cl
2 Jul 2026
Safety

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine

DGX agent

arXiv:2607.00576v1 Announce Type: new Abstract: Multi-image content has become an increasingly prevalent form of visual communication in social media, giving rise to a new safety issue, multi-image im

safetyarxiv-cs-cl
2 Jul 2026
Safety

Selective Test-Time Debiasing for CLIP via Reward Gating

DGX agent

arXiv:2607.00423v1 Announce Type: new Abstract: Vision language models (VLMs) demonstrate strong zero-shot performance, but often perpetuate social stereotypes in person-centric queries, yielding skew

safetyarxiv-cs-cl
2 Jul 2026
Research

SemEval-2026 Task 9: Detecting Multilingual, Multicultural and Multievent Online Polarization

DGX agent

arXiv:2604.06817v2 Announce Type: replace Abstract: We present SemEval-2026 Task 9, a shared task on online polarization detection, covering 22 languages and comprising over 110K annotated instances.

researcharxiv-cs-cl
2 Jul 2026
Agents

SlowBA: An efficiency backdoor attack towards VLM-based GUI agents

DGX agent

arXiv:2603.08316v3 Announce Type: replace-cross Abstract: Modern vision-language-model (VLM) based graphical user interface (GUI) agents are expected not only to execute actions accurately but also to

agentsarxiv-cs-cl
2 Jul 2026
Safety

Speech Playground: An Interactive Tool for Speech Analysis and Comparison

DGX agent

arXiv:2607.00418v1 Announce Type: new Abstract: This paper presents Speech Playground, an interactive speech visualization and comparison tool. While existing tools such as Praat are excellent, it can

safetyarxiv-cs-cl
2 Jul 2026
Model Releases

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning

DGX agent

arXiv:2607.00465v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) rely extensively on Visual Instruction Tuning (VIT) to elicit their multimodal reasoning capabilities. However, w

model-releasesarxiv-cs-cl
2 Jul 2026
Research

Structural Pattern Mining in Inka Khipus: Unsupervised Clustering, Provenance Classification, and a Computational Validation of the Santa Valley Match

DGX agent

arXiv:2607.00185v1 Announce Type: new Abstract: Khipus--knotted cord devices--were the primary recording medium of the Inka Empire (c. 1400-1532 CE), yet their system remains undeciphered. We present

researcharxiv-cs-cl
2 Jul 2026
← Previous
1…3435363738…161
Next →