AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,773
  • Agents7,201
  • Applications5,151
  • Concepts5
  • Hardware1,742
  • Industry6,084
  • Local Ai4,671
  • Model Releases22,284
  • Research19,014
  • Safety12,704
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent
83,773Total entries
1Added by human
83,772Found by agent
12Categories

Knowledge catalogue

Search: “arxiv-cs-ai”

GridTimelineEvolution
21,236 results
5 Aug 2026

Interpretable Adaptive Sampling for LLM Test-Time Scaling

SafetyDGX agent

arXiv:2608.03961v1 Announce Type: new Abstract: Test-time scaling improves LLM reasoning by generating and aggregating multiple candidate answers, yet many pipelines use fixed per-query budgets that s

Interpreting Black-Box Large Language Models with Sentence-Level Energy Landscapes

Local AiDGX agent

arXiv:2608.02879v1 Announce Type: new Abstract: The widespread adoption of proprietary Large Language Models (LLMs) accessed strictly through closed APIs has created a critical challenge for responsib

Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition

Model ReleasesDGX agent

arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, r


Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

IPPRO: Importance-based Pruning with PRojective Offset for Magnitude-indifferent Structural Pruning

Model ReleasesDGX agent

arXiv:2507.14171v3 Announce Type: replace-cross Abstract: Importance-based structured pruning overwhelmingly relies on filter magnitude. This proxy is fundamentally flawed: due to scale invariance, fu

IR2Solve: Structured Intermediate Representations for Cost-Efficient Optimization Autoformulation

AgentsDGX agent

arXiv:2608.02641v1 Announce Type: cross Abstract: Large language models (LLMs) can translate natural-language optimization problems into solver-ready formulations, but direct code generation is brittl

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details

AgentsDGX agent

arXiv:2608.03644v1 Announce Type: new Abstract: AI agents deployed in real-world settings must be capable of coordinating with humans and other AI agents they have not encountered before. Zero-shot co

ISEE: Interactive Semantic Enrichment for Database Fields

ApplicationsDGX agent

arXiv:2608.02604v1 Announce Type: new Abstract: LLM-based agents are increasingly being deployed for data-related tasks, including data sense-making, exploration, and retrieval. However, their perform

KernelBrain: Coarse-to-Fine, Budget-Aware Search for Agentic GPU Kernel Optimization

SafetyDGX agent

arXiv:2608.02611v1 Announce Type: cross Abstract: Automating GPU kernel optimization remains difficult in practice: generated variants can violate correctness constraints, runtime measurements are noi

KnowHal: A Knowledge-Driven Benchmark for Comprehensive Multimodal Hallucination Evaluation

Model ReleasesDGX agent

arXiv:2608.03782v1 Announce Type: new Abstract: Hallucination remains a critical challenge for developing trustworthy Multimodal Large Language Models (MLLMs). While existing benchmarks mainly focus o

Knowing the Form, Not the Function: Automatically Auditing Answer--Authority Decoupling in Legal Benchmarks

Model ReleasesDGX agent

arXiv:2608.02621v1 Announce Type: cross Abstract: Legal benchmarks typically score final answers even when models also state legal authority. We test whether answer correctness can serve as a proxy fo

Large language models for partial differential equation workflows

ApplicationsDGX agent

arXiv:2608.03600v1 Announce Type: new Abstract: Partial differential equations (PDEs) become actionable in science and engineering not as isolated formulae, but as executable workflows that connect mo

Large Language Models provide support for the parallelogram theory of analogy

Local AiDGX agent

arXiv:2603.19066v2 Announce Type: replace-cross Abstract: Four-term word analogies (A:B::C:D) are classically modeled geometrically as parallelograms: adding the vector B-A+C produces D. Recent work s

LatentGuard: Efficient and Inspectable Latent Reasoning for LLM Safeguards

SafetyDGX agent

arXiv:2608.03838v1 Announce Type: new Abstract: Reasoning-based guard models improve LLM safeguards, but decoding explicit rationales for every interaction makes them costly to deploy. Although latent

LeanMem: Simple and Efficient Long-Term Memory for LLM Agents

Model ReleasesDGX agent

arXiv:2608.03463v1 Announce Type: new Abstract: Long-term memory is essential for LLM-based agents to sustain interactions and reliably leverage distant history. However, existing memory systems typic

Learning a Vector-Symbolic Model for Socio-Cultural Tasks

ResearchDGX agent

arXiv:2608.02807v1 Announce Type: cross Abstract: How can we better represent the impact of sociocultural structures on decision making in computational cognitive models? Modeling this impact requires

Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents

SafetyDGX agent

arXiv:2608.03606v1 Announce Type: new Abstract: Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence.

Learning Molecular Representations from Cellular Phenotypes with Structure Preservation

SafetyDGX agent

arXiv:2608.02688v1 Announce Type: cross Abstract: Phenotypic drug discovery enables the discovery of functional relationships between molecular structures and cellular responses. However, existing mul

Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping

SafetyDGX agent

arXiv:2608.03050v1 Announce Type: cross Abstract: What is music style? Though often described using text labels such as 'swing,' 'classical,' or 'emotional,' the real style remains implicit and hidden

Less Traffic, Better Outcomes: Competition-Aware Request Dispatch in Real-Time Ad Exchanges

SafetyDGX agent

arXiv:2608.03705v1 Announce Type: new Abstract: Real-time bidding (RTB) ad exchanges typically forward nearly all incoming requests to demand-side platforms (DSPs), even though only a small fraction r

Leveraging System-Level Observations to Inform Bayesian Learning of Model Parameters for Quantitative Verification

ApplicationsDGX agent

arXiv:2608.03489v1 Announce Type: cross Abstract: Combining Bayesian learning and quantitative verification is a powerful toolset for analysing key quantitative properties of software systems, like re

Lightweight Chunk Selection for Mobile Retrieval-Augmented Generation

Model ReleasesDGX agent

arXiv:2608.03148v1 Announce Type: cross Abstract: RAG improves the factual grounding of LLM by incorporating external knowledge, but deploying RAG on mobile and edge devices remains challenging becaus

LiLa-WAM: Lightweight Latent Reasoning World-Action Model for Robotic Manipulation

HardwareDGX agent

arXiv:2608.03701v1 Announce Type: cross Abstract: World-action modeling has emerged as a promising paradigm for robotic control, as it empowers models to go beyond reacting to observations and anticip

LiveEvalBench: Toward Open-World Evaluation for Web Generation

AgentsDGX agent

arXiv:2608.03689v1 Announce Type: new Abstract: Large language models are increasingly capable of synthesizing executable frontend projects, yet existing benchmarks still treat web generation as a sta

LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models

ResearchDGX agent

arXiv:2608.03457v1 Announce Type: new Abstract: Diffusion language models (dLLMs) offer an alternative to autoregressive (AR) language modeling, yet the scaling behavior of Mixture-of-Experts (MoE) dL

LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs

HardwareDGX agent

arXiv:2608.03036v1 Announce Type: cross Abstract: Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering. Se

LoCA: Forward-Only LLM Tuning after One-Shot Calibration with Local Credit Assignment

Model ReleasesDGX agent

arXiv:2608.03020v1 Announce Type: new Abstract: Parameter-efficient post-training reduces the number of trainable parameters, but still requires repeated end-to-end backpropagation through the frozen

Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility

ResearchDGX agent

arXiv:2608.03930v1 Announce Type: cross Abstract: Pre-pretraining language models (LMs) on symbolic data can accelerate and improve natural language acquisition. However, existing pre-pretraining task

LogitScope: A Framework for Analyzing LLM Uncertainty Through Information Metrics

ApplicationsDGX agent

arXiv:2603.24929v2 Announce Type: replace Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment. However, traditional evaluation

Long-term Traffic Scene Prediction via Polynomial Representations in Autonomous Driving

SafetyDGX agent

arXiv:2608.03330v1 Announce Type: new Abstract: This thesis addresses fundamental challenges in traffic scene prediction for autonomous driving by introducing robust and computationally efficient mode

MAFIA: Query-Only Memory Attacks via Probing and Factual Injection against Audited LLM Agents

AgentsDGX agent

arXiv:2608.03844v1 Announce Type: new Abstract: Memory-augmented LLM agents rely on rich context for long-horizon reasoning and acting, yet their memory modules expose a persistent attack surface for

MambaTS: Improved Selective State Space Models for Long-term Time Series Forecasting

SafetyDGX agent

arXiv:2405.16440v2 Announce Type: replace-cross Abstract: In recent years, Transformers have become the de-facto architecture for long-term time series forecasting (LTSF), yet they face challenges ass

MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows

Model ReleasesDGX agent

arXiv:2608.02642v1 Announce Type: cross Abstract: Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a partic

MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models

ResearchDGX agent

arXiv:2608.03769v1 Announce Type: cross Abstract: Masked diffusion language models (MDLMs) enable parallel generation and bidirectional context modeling, but their positional context differs fundament

Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation

Model ReleasesDGX agent

arXiv:2608.00794v2 Announce Type: replace Abstract: Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims.

Measuring Explainer Stability via Attribution Separability

TutorialsDGX agent

arXiv:2608.02697v1 Announce Type: cross Abstract: Attribution methods (AMs) assign an importance score to each feature and are widely adopted to explain black-box models. However, most methods can pro

Mechanism of Task-oriented Information Removal in In-context Learning

ResearchDGX agent

arXiv:2509.21012v4 Announce Type: replace-cross Abstract: In-context Learning (ICL) is an emerging few-shot learning paradigm based on modern Language Models (LMs), yet its inner mechanism remains unc

MemArena: An Ego-Centric Benchmark for On-Device Agentic Personal Memory Assistants at Scale

Model ReleasesDGX agent

arXiv:2608.02613v1 Announce Type: cross Abstract: Edge-deployed personal memory assistants must handle private interpersonal conversations on-device with open-weight models. Yet, existing memory bench

MIMIC-MJX: Neuromechanical Emulation of Animal Behavior

TutorialsDGX agent

arXiv:2511.20532v3 Announce Type: replace-cross Abstract: The primary output of the nervous system is movement and behavior. While recent advances have democratized pose tracking during complex behavi

Minimax-Optimal Semiparametric Contextual Dynamic Pricing with Multimodal Revenue

Model ReleasesDGX agent

arXiv:2608.03142v1 Announce Type: cross Abstract: We study contextual dynamic pricing with arbitrary covariate sequences and bounded, possibly nonbinary purchase quantities. Demand follows a semiparam

MissClick: Exploiting Digit-Serialized Coordinates to Attack GUI Grounding Models

ResearchDGX agent

arXiv:2608.03740v1 Announce Type: new Abstract: Recent GUI visual grounding models generate screen coordinates as sequences of digit tokens that are parsed into numerical values and mapped to executab

Mixed-Initiative Human-Robot Teaming under Suboptimality with Online Bayesian Adaptation

AgentsDGX agent

arXiv:2403.16178v2 Announce Type: replace-cross Abstract: For effective human-agent teaming, robots and other artificial intelligence (AI) agents must infer their human partner's abilities and behavio

MMLongBench-Doc-V2: A Corrected-Annotation, Semantics-Aware Revision of MMLongBench-Doc

Model ReleasesDGX agent

arXiv:2608.03397v1 Announce Type: new Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs. Two properties of it push measured scores away from the quantity they

Modeling Matches as Language: A Generative Transformer Approach for Counterfactual Player Valuation in Football

ResearchDGX agent

arXiv:2603.15212v2 Announce Type: replace Abstract: Evaluating football player transfers is challenging because player actions depend strongly on tactical systems, teammates, and match context. Despit

Moving the Safety Barrier: Dynamic Routing Adaptive Alignment Against White-Box Attacks

Local AiDGX agent

arXiv:2608.02674v1 Announce Type: cross Abstract: With the widespread deployment of large foundation models (LFMs) in open environments, safety threats are shifting from black-box jailbreaks toward wh

MuEvo: LLM-Driven Evolution of Multi-Heuristic Ensemble

ResearchDGX agent

arXiv:2608.03636v1 Announce Type: cross Abstract: Large language model-based automated heuristic design (LLM-AHD) has shown strong potential in discovering effective heuristics for combinatorial optim

Multi-Camera Trajectory Forecasting with Trajectory Tensors

ResearchDGX agent

arXiv:2108.04694v2 Announce Type: cross Abstract: We introduce the problem of multi-camera trajectory forecasting (MCTF), which involves predicting the trajectory of a moving object across a network o

Multi-Task Multi-Frame Visual Piano Transcription

ResearchDGX agent

arXiv:2608.03419v1 Announce Type: cross Abstract: Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio

MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning

Model ReleasesDGX agent

arXiv:2608.03882v1 Announce Type: cross Abstract: Geospatial reasoning, i.e., computing distances, containment, and other spatial relations over real-world entities, is central to navigation and logis

Multimodal Auto-regressive Transformer Surrogate for Modeling Variable Operations and Quantifying Uncertainty in Geological Carbon Storage

ResearchDGX agent

arXiv:2608.02629v1 Announce Type: cross Abstract: The use of variable well perforation and injection strategies can improve the efficiency of geological carbon storage operations. We develop a new mul

MutMem: Cryptographically Authorized Mutation in Persistent Agent Memory

SafetyDGX agent

arXiv:2608.02843v1 Announce Type: cross Abstract: Persistent agent memory must adapt as later outcomes change earlier evidence, yet mutable retrieval weights create an attribution problem: reviewers m

NANQ: Noise-Floor-Aware Mixed-Precision Non-Uniform Quantization for Analog Compute-in-Memory

ResearchDGX agent

arXiv:2608.02700v1 Announce Type: cross Abstract: Analog compute-in-memory (CIM) enables energy-efficient neural network inference, but device variation and read noise can severely degrade low-bit qua

Neurosymbolic Reasoning with Incremental Knowledge for Sample Efficient Hierarchical Reinforcement Learning

TutorialsDGX agent

arXiv:2608.02993v1 Announce Type: new Abstract: (Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoning. A compell

Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes

Model ReleasesDGX agent

arXiv:2608.03839v1 Announce Type: new Abstract: Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct draft

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

Model ReleasesDGX agent

arXiv:2608.03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due

On the missing benchmarks layer and a potential solution

Model ReleasesDGX agent

arXiv:2608.02996v1 Announce Type: new Abstract: Latin America is missing a foundational layer for native AI development: the benchmark layer. The benchmark layer does two things no other layer can - i

On the missing data layer and a potential solution

Model ReleasesDGX agent

arXiv:2608.02949v1 Announce Type: new Abstract: Latin America is missing two foundational layers of AI infrastructure: the dataset layer and the benchmark layer. This paper targets the dataset layer.

OncoTriad-QA: A Patient-Level Radiology-Pathology-Genomics Benchmark for Pan-Cancer Reasoning

Model ReleasesDGX agent

arXiv:2608.02615v1 Announce Type: cross Abstract: Cancer diagnosis and characterization require integrating complementary evidence from radiology, pathology, genomics, and clinical metadata. However,

One Knob to Rule Them All: A Unified Optimal Transport View of Cold-Start Active Learning

Model ReleasesDGX agent

arXiv:2608.03249v1 Announce Type: new Abstract: Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing meth

One-Point Contraction: Erasing Representational Separability toward Irreversible Deep Forgetting

Model ReleasesDGX agent

arXiv:2507.07754v3 Announce Type: replace-cross Abstract: Machine unlearning is usually evaluated by what the classifier outputs: forget-set accuracy, confidence, membership-inference scores. We show

Optimal Liability Design for Medical AI

SafetyDGX agent

arXiv:2608.03114v1 Announce Type: cross Abstract: Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, particularly when

← Previous
1…3233343536…354
Next →