AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-cl”

GridTimelineEvolution
7,646 results
4 Aug 2026

Does Accuracy Equal Evidence? Reasoning Faithfulness under KV Cache Compression

ResearchDGX agent

arXiv:2608.01631v1 Announce Type: new Abstract: KV cache compression is commonly evaluated by final-answer accuracy, implicitly assuming that preserving the answer also preserves the reasoning that su

Does Machine 'know' interpersonal pragmatics? Evidence from MARBERT's learning of emoji pragmatics in Arabic digital discourse

TutorialsDGX agent

arXiv:2608.01174v1 Announce Type: new Abstract: This study examines Transformer-based models' ability to learn emoji pragmatics in Arabic digital discourse (ADD), providing evidence from MARBERT's beh

Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result

ApplicationsDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2608.01559v1 Announce Type: cross Abstract: Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the

Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

Model ReleasesDGX agent

arXiv:2608.02235v1 Announce Type: new Abstract: Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However

Don't Judge a Book by its Cover: Testing LLMs' Robustness Under Logical Obfuscation

Model ReleasesDGX agent

arXiv:2602.01132v2 Announce Type: replace Abstract: Tasks such as solving arithmetic equations, evaluating truth tables, and completing syllogisms are handled well by large language models (LLMs) in t

Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale

ApplicationsDGX agent

arXiv:2608.01050v1 Announce Type: cross Abstract: Production LLM agents that select from large skill libraries face a limitation that semantic relevance alone cannot resolve: a skill may match a user'

Douyin Multimodal Embedding Model Technical Report

ApplicationsDGX agent

arXiv:2608.02148v1 Announce Type: cross Abstract: Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search

EHR2Path: Comprehensive Pathway-Level Modeling of Longitudinal Patient Trajectories from Multimodal Electronic Health Records

ResearchDGX agent

arXiv:2506.04831v3 Announce Type: replace-cross Abstract: Forecasting how a patient's condition is likely to evolve, including possible deterioration, recovery, treatment needs, and care transitions,

Enhancing Large Language Model Reasoning with Reward Models: An Analytical Survey

ResearchDGX agent

arXiv:2510.01925v3 Announce Type: replace Abstract: Reward models (RMs) play a critical role in enhancing the reasoning performance of LLMs. For example, they can provide training signals to finetune

Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning

TutorialsDGX agent

arXiv:2608.00994v1 Announce Type: cross Abstract: Zero-shot image captioning aims to generate image descriptions without annotated image-text pairs. Recent approaches exploit text-to-image models to s

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs

Model ReleasesDGX agent

arXiv:2608.01979v1 Announce Type: cross Abstract: Visual token pruning reduces the inference cost of multimodal large language models, but a fixed token ratio is poorly matched to text-rich inputs. In

Evaluating VLMs on Multimodal Aristotelian Persuasion Tasks

Model ReleasesDGX agent

arXiv:2608.01238v1 Announce Type: new Abstract: Vision Language Models (VLMs) have demonstrated exceptional performance across various tasks. However, they have not yet been thoroughly evaluated on mo

EviSD: Evidence-Conditioned Self-Distillation for Search-Augmented Agents

Local AiDGX agent

arXiv:2608.01359v1 Announce Type: new Abstract: Outcome-based reinforcement learning enables search-augmented language agents to learn from verifiable final answers, but its trajectory-level credit ca

Exemplars in Disguise: Pure Exemplar Models Mimic Abstraction-First Learning

TutorialsDGX agent

arXiv:2608.00821v1 Announce Type: new Abstract: Whether idiosyncratic, item-specific knowledge is learned before abstract class-level generalizations, or vice versa, is a central question in language

Expert-Choice Routing Enables Adaptive Computation in Diffusion Language Models

SafetyDGX agent

arXiv:2604.01622v2 Announce Type: replace-cross Abstract: Diffusion language models (DLMs) enable parallel, non-autoregressive text generation, yet existing DLM mixture-of-experts (MoE) models inherit

Exploiting Intrinsic Duality for Multi-Hop Question Generation

SafetyDGX agent

arXiv:2608.00712v1 Announce Type: new Abstract: Multi hop question generation (MQG) aims to generate questions from multiple given documents and target answers, whereas question answering (QA) focuses

Exploring More to Solve More: Boosting Diversity in Text Diffusion Models via Entropy-Based Guidance

ResearchDGX agent

arXiv:2608.00024v1 Announce Type: new Abstract: Although diffusion models have revolutionized continuous domains like image synthesis through high quality generations and controllable guidance mechani

Fast and Accurate Quotation Attribution in Literary Texts

Model ReleasesDGX agent

arXiv:2608.02359v1 Announce Type: new Abstract: Attributing quotations to their speakers in literary texts remains an open challenge. Standard methods, which independently predict a speaker mention fo

Fenced Citation-Context Retrieval for Case Law: Temporal Leakage and Degree Control Across Two Jurisdictions

ResearchDGX agent

arXiv:2607.17142v3 Announce Type: replace-cross Abstract: Prior case retrieval (PCR) aims to identify the precedent cases relevant to the facts of a query case. Incoming citation context, the text wit

Few-Shot Biomedical Relation Extraction with Large Language Models: A Viable Alternative to Supervised Learning?

ResearchDGX agent

arXiv:2606.15412v2 Announce Type: replace Abstract: Biomedical relation extraction (BioRE) is a key step in transforming biomedical literature into structured knowledge. Most existing approaches rely

FinHardBench: Can LLMs Generate Latency-Aware Hardware for Financial Computing?

Model ReleasesDGX agent

arXiv:2608.00909v1 Announce Type: new Abstract: Can large language models generate not just correct, but fast hardware? This paper investigates the question in financial FPGA design, where 5-10 nanose

Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?

Model ReleasesDGX agent

arXiv:2608.01704v1 Announce Type: cross Abstract: A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both boun

From Chains to Trees: Parent-Conditioned Drafting for Semi-Autoregressive Speculative Decoding

ResearchDGX agent

arXiv:2608.02123v1 Announce Type: new Abstract: Speculative decoding accelerates LLM inference only when drafted continuations survive target-model verification. Semi-autoregressive drafters such as D

From Direction to Magnitude: How Multimodal Instruction-Tuning Reorganizes the Geometric Encoding of Identity-Specifying Prompts in Transformer Hidden States

Model ReleasesDGX agent

arXiv:2607.09842v2 Announce Type: replace-cross Abstract: We investigate whether identity-specifying system prompts produce statistically distinguishable geometric fingerprints in the hidden-state tra

From We to Me: Theory Informed Narrative Shift with Abductive Reasoning

Model ReleasesDGX agent

arXiv:2603.03320v2 Announce Type: replace Abstract: Effective communication often relies on aligning a message with an audience's narrative and worldview. Narrative shift involves transforming text to

Gaokerena: A Small Persian Medical Language Model Family

Model ReleasesDGX agent

arXiv:2608.00932v1 Announce Type: new Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused

Geometry-Guided Layerwise FFN Width Allocation in Transformers

ResearchDGX agent

arXiv:2608.02064v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) account for a large fraction of Transformer parameters, yet their hidden width is usually constant across depth. We ask w

Global Optimization and Inference-Time Region Grafting for Agentic Workflows

Model ReleasesDGX agent

arXiv:2608.02353v1 Announce Type: new Abstract: Recent advances in agentic workflow optimization automate workflow design through task-specific workflow search or input-conditioned architecture select

GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning

ResearchDGX agent

arXiv:2608.02585v1 Announce Type: cross Abstract: Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models

ResearchDGX agent

arXiv:2608.02124v1 Announce Type: cross Abstract: Vision-language models (VLMs) remain unreliable when predictions require fine-grained visual evidence. We identify a previously overlooked cause: spec

HarnessCompass: Guiding Automatic Harness Evolution toward Generalizable and Effective Agent Harnesses

Model ReleasesDGX agent

arXiv:2608.01918v1 Announce Type: cross Abstract: Harness design plays a critical role in agent performance by shaping how large language models (LLMs) perceive, reason over, and act within executable

HetRoute Heterogeneous and Cost-aware Collaborative Routing Framework for Distributed Edge MoE Inference

HardwareDGX agent

arXiv:2608.00577v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) models have become a dominant architecture for large-scale AI services, yet deploying them over geo-distributed heterogeneous

Hierarchical Pre-Training of Vision Encoders with Large Language Model

SafetyDGX agent

arXiv:2604.00086v2 Announce Type: replace-cross Abstract: The field of computer vision has experienced significant advancements through scalable vision encoders and multimodal pre-training frameworks.

HopRefusalBench: Diagnosing Refusal Failures in Search-Augmented Agents for Multi-Hop Reasoning

Model ReleasesDGX agent

arXiv:2608.01358v1 Announce Type: new Abstract: Search-augmented large language model agents are increasingly capable of solving knowledge-intensive tasks, but their behavior when a multi-hop question

Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese

SafetyDGX agent

arXiv:2608.01629v1 Announce Type: new Abstract: Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speak

Hylog: A Hybrid Approach to Logging Text Production in Non-alphabetic Scripts

ApplicationsDGX agent

arXiv:2601.17753v2 Announce Type: replace Abstract: Research keyloggers are essential for cognitive studies of text production, yet most fail to capture the on-screen transformations performed by Inpu

IACM-RL: Intent-Aware Context Management and Reinforcement Learning for Complex Tool Invocation under Dynamic Intent Fluctuations

SafetyDGX agent

arXiv:2608.02110v1 Announce Type: new Abstract: Executing long-horizon tool invocations in real-world environments is severely challenged by dynamic user intent noise. Existing methods attempt robustn

Illuminating Visual Identity in Universal Multimodal Embeddings

Model ReleasesDGX agent

arXiv:2608.01794v1 Announce Type: cross Abstract: Universal Multimodal Embeddings (UMEs) aim to unify various modalities and tasks into a shared representation space. In recent years, this field has w

Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation

SafetyDGX agent

arXiv:2608.02087v1 Announce Type: cross Abstract: Post-training Large Language Models (LLMs) with Reinforcement Learning (RL) has become an important tool for improving model capabilities, but the LLM

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving

AgentsDGX agent

arXiv:2512.10739v3 Announce Type: replace Abstract: Large Reasoning Models (LRMs) have expanded the mathematical reasoning frontier through Chain-of-Thought (CoT) techniques and Reinforcement Learning

Interpretable Recognition of Cognitive Distortions in Natural Language Texts

ResearchDGX agent

arXiv:2511.05969v2 Announce Type: replace Abstract: We propose a new approach to multi-factor classification of natural language texts based on weighted structured patterns such as N-grams, taking int

Just on Time: Token-Level Early Stopping for Diffusion Language Models

ResearchDGX agent

arXiv:2602.11133v2 Announce Type: replace-cross Abstract: Diffusion language models generate text through iterative refinement, a process that is often computationally inefficient because many tokens

LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering

Model ReleasesDGX agent

arXiv:2604.03532v2 Announce Type: replace Abstract: Large language models (LLMs) show strong multilingual capabilities, yet reliably controlling the language of their outputs remains difficult. Repres

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

Local AiDGX agent

arXiv:2608.01395v1 Announce Type: new Abstract: We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU la

Leak It: A Probabilistic Approach to Training-Data Extraction from Black-Box Language Models

ResearchDGX agent

arXiv:2608.00144v1 Announce Type: cross Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind bas

Learning What to Remember: Test-Time Training via Context Distillation

Model ReleasesDGX agent

arXiv:2608.01672v1 Announce Type: new Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later. Test

Length Penalties Make Chain-of-Thought Less Monitorable

ResearchDGX agent

arXiv:2607.09786v3 Announce Type: replace-cross Abstract: To curb overthinking and reduce inference costs, researchers now train reasoning models with penalties on chain of thought length. We find tha

Leveraging Synthetic Data for Question Answering with Multilingual LLMs in the Agricultural Domain

ResearchDGX agent

arXiv:2507.16974v3 Announce Type: replace Abstract: Enabling farmers to access accurate agriculture-related information in their native languages in a timely manner is crucial for the success of the a

LiveMem: Maintaining Memory State Continuity in Long-Running LLM Inference

Model ReleasesDGX agent

arXiv:2608.02515v1 Announce Type: new Abstract: Long-running assistants and agents consume interaction streams that eventually outgrow the context. Existing context retention, summarization, and retri

LLM generation novelty through the lens of semantic similarity

ResearchDGX agent

arXiv:2510.27313v3 Announce Type: replace-cross Abstract: Generation novelty is a key indicator of an LLM's ability to generalize, yet measuring it against full pretraining corpora is computationally

LLM-OSDA: An Optimal-Stopping Dynamic Auction for Native Advertising in Multi-Turn LLM Conversations

SafetyDGX agent

arXiv:2608.00123v1 Announce Type: new Abstract: LLM-native advertising embeds sponsored content directly into model-generated responses, shifting the unit of sale from a fixed slot to a moment within

LM-mixup: Text Data Augmentation via Language Model based Mixup

SafetyDGX agent

arXiv:2510.20449v2 Announce Type: replace Abstract: Instruction tuning is crucial for aligning Large Language Models (LLMs), yet the quality of instruction-following data varies significantly. While h

Loanword or Switch? The Annotation Boundary, Not the Model, Drives Kazakh-Russian Code-Switching Identification

ResearchDGX agent

arXiv:2608.00581v1 Announce Type: new Abstract: Off-the-shelf LID and letter heuristics over-label Kazakh-Russian social text as mixed: Russian loanwords inside Kazakh look like code-switching under a

Long-Horizon Embodied Decision-Making via Multimodal Memory Compression

Model ReleasesDGX agent

arXiv:2608.01456v1 Announce Type: cross Abstract: Agents are increasingly expected to act not only as task executors, but also as decision-makers on behalf of human users. This shift requires agents t

LongCat Sparse Attention: Taming the Lightning via Streaming-aware Hierarchical Cross-Layer Indexing

Model ReleasesDGX agent

arXiv:2608.01662v1 Announce Type: cross Abstract: DeepSeek Sparse Attention (DSA) enables efficient long-context modeling through its Lightning Indexer. However, practical deployment remains constrain

LongChart VQA: A Comprehensive Benchmark for MLLMs with Complex Multi-Chart Reasoning

Model ReleasesDGX agent

arXiv:2608.01328v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) are rapidly evolving with expanded context windows and stronger reasoning capabilities, enabling multi-chart un

Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

SafetyDGX agent

arXiv:2608.01953v1 Announce Type: new Abstract: On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference

LoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding Agent

Model ReleasesDGX agent

arXiv:2608.00267v1 Announce Type: cross Abstract: Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon soft

MAPLE: Metadata Augmented Private Language Evolution

Model ReleasesDGX agent

arXiv:2603.19258v2 Announce Type: replace Abstract: Differentially private (DP) fine-tuning of large language models (LLMs) requires massive compute and full model access, which rules out state-of-the

MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs

Model ReleasesDGX agent

arXiv:2608.02520v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than

← Previous
1…89101112…128
Next →