AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries86,542
  • Agents7,406
  • Applications5,305
  • Concepts5
  • Hardware1,791
  • Industry6,129
  • Local Ai4,837
  • Model Releases23,234
  • Research19,717
  • Safety13,103
  • Syntheses17
  • Tools1,670
  • Tutorials3,328

Source
HumanDGX agent

Content type
86,542Total entries
1Added by human
86,541Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
50,764 results
Model Releases

AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities

DGX agent

arXiv:2607.13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluat

model-releasesarxiv-cs-ai
16 Jul 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

HEDGEHOG: Hierarchical Evaluation of Drug Generators Through Rigorous Filtration

DGX agent

arXiv:2607.13155v1 Announce Type: new Abstract: Generative molecular models can support early drug discovery by proposing new candidate compounds de novo. In practice, useful candidates must balance t

model-releasesarxiv-cs-lg
16 Jul 2026
Model Releases

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

DGX agent

arXiv:2607.13069v1 Announce Type: new Abstract: Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We intr

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

MedRealMM: A Real-World Multimodal Benchmark for Chinese Online Medical Consultation

DGX agent

arXiv:2607.09142v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in online medical consultation, yet existing benchmarks remain poorly aligned with real clini

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

SPINE: Bridging the Cyber-Physical Gap with Agentic AI

DGX agent

arXiv:2607.13049v1 Announce Type: new Abstract: Foundation models have given robots a sophisticated brain for complex decision-making, yet deploying that intelligence into a physical platform still de

model-releasesarxiv-cs-ai
16 Jul 2026
Safety

Uncertainty-Aware Sequential Decision Rules for Event-Triggered LLM Invocation in Streaming Systems

DGX agent

arXiv:2607.13048v1 Announce Type: cross Abstract: Streaming inference pipelines increasingly pair lightweight fast models with Large Language Models (LLMs) that provide rich semantic understanding at

safetyarxiv-cs-ai
16 Jul 2026
Model Releases

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs

DGX agent

arXiv:2602.17659v2 Announce Type: replace Abstract: Vision-Language-Action models (VLAs) promise to ground language instructions in robot control, yet in practice often fail to faithfully follow langu

model-releasesarxiv-cs-cv
16 Jul 2026
Model Releases

1D-Bench: A Benchmark for Iterative UI Code Generation with Visual Feedback in Real-World

DGX agent

arXiv:2602.18548v2 Announce Type: replace-cross Abstract: Design-to-code translates high-fidelity UI designs into executable front-end implementations, but progress remains hard to compare due to inco

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

A JoLT for the KV Cache: Near-Lossless KV Cache Compression via Joint Tucker and JL-Residual Allocation for LLMs

DGX agent

arXiv:2607.12550v1 Announce Type: cross Abstract: The key-value (KV) cache has become the dominant memory cost of transformer inference. It grows with batch size, context length, and depth, and at lon

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

CANDI: Contextual Alignment for Niche Domains Question Answering

DGX agent

arXiv:2607.11891v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) in specialized domains like medical diagnostics and financial advisory necessitates evaluating capabili

model-releasesarxiv-cs-ai
15 Jul 2026
Model Releases

Controllable Generation of Diverse Dermatological Imagery for Fair and Efficient Malignancy Classification

DGX agent

arXiv:2607.12987v1 Announce Type: new Abstract: Accurate dermatological diagnosis naturally necessitates equitable performance across diverse populations, yet a systematic lack of expertly annotated i

model-releasesarxiv-cs-cv
15 Jul 2026
Model Releases

Label-Decoupled Style Augmentation for Domain Generalization in Multi-Label Remote Sensing Scene Classification

DGX agent

arXiv:2607.12704v1 Announce Type: new Abstract: Multi-label classification assigns several co-occurring labels to each aerial scene, yet deployed models often encounter data distributions different fr

model-releasesarxiv-cs-cv
15 Jul 2026
Tutorials

Language Identification with Succinct Machine-Independent Traces

DGX agent

arXiv:2607.12443v1 Announce Type: new Abstract: Motivated by the power of large language models, there has been renewed interest in the Gold-Angluin model of language identification in the limit, with

tutorialsarxiv-cs-cl
15 Jul 2026
Research

LLM Judges Can Be Too Generous When There Is No Reference Answer

DGX agent

arXiv:2607.12885v1 Announce Type: new Abstract: LLM judges are increasingly being used to evaluate open-ended model responses, often in no-reference settings where a ground-truth answer is unavailable

researcharxiv-cs-cl
15 Jul 2026
Local Ai

More Than Where You Are: Learning Semantics, Structure, and Geometry from Cross-View Localization

DGX agent

arXiv:2607.12429v1 Announce Type: new Abstract: Consistent cross-view understanding under extreme viewpoint changes is essential for spatial intelligence, as it enables models to recognize the same sc

local-aiarxiv-cs-cv
15 Jul 2026
Model Releases

Predict the Retrieval! Test time adaptation for Retrieval Augmented Generation

DGX agent

arXiv:2601.11443v3 Announce Type: replace Abstract: Retrieval-Augmented Generation (RAG) has emerged as a powerful approach for enhancing large language models' question-answering capabilities through

model-releasesarxiv-cs-cl
15 Jul 2026
Model Releases

RecRec: Recursive Refinement for Sequential Recommendation

DGX agent

arXiv:2607.10541v2 Announce Type: replace-cross Abstract: Sequential recommender systems typically infer user preferences through single-pass encoding of interaction histories without iterative refine

model-releasesarxiv-cs-lg
15 Jul 2026
Model Releases

Self-Evolving In-Context Learning for Direct Pilot-to-Beamformer Design in MU-MISO Systems

DGX agent

arXiv:2607.11970v1 Announce Type: cross Abstract: We develop an enhanced in-context learning (ICL) framework to improve the performance of pilot-based beamforming in multi-user multiple-input single-o

model-releasesarxiv-cs-ai
15 Jul 2026
Research

The Spectrum Is Not Enough: When Context Helps Time-Series Forecasting

DGX agent

arXiv:2607.13006v1 Announce Type: new Abstract: A growing family of indices scores how predictable a series is from its spectrum. Practitioners increasingly read these scores as answering a different

researcharxiv-cs-lg
15 Jul 2026
Model Releases

TRACE: An Operational Reasoning Schema for Auditable Agentic Commitments

DGX agent

arXiv:2607.12480v1 Announce Type: new Abstract: This paper defines TRACE (Typed Reasoning And Commitment Evidence): a typed, versioned schema for recording reasoning traces, a reference procedure for

model-releasesarxiv-cs-ai
15 Jul 2026
Local Ai

Collate: Collaborative Neural Network Learning for Latency-Critical Edge Systems

DGX agent

arXiv:2607.08013v1 Announce Type: new Abstract: Federated Learning (FL) empowers multiple clients to collaboratively learn a model, enlarging the training data of each client for high accuracy while p

local-aiarxiv-cs-lg
10 Jul 2026
Model Releases

Dive Into the Implicit Biases of Low-rank Vision-language Alignment

DGX agent

arXiv:2607.08194v1 Announce Type: new Abstract: Vision-language alignment, the stage that bridges pretrained vision encoders and large language models, is widely treated as a form of pretraining requi

model-releasesarxiv-cs-cv
10 Jul 2026
Applications

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

DGX agent

arXiv:2607.08009v1 Announce Type: new Abstract: We introduce a Bloom-aligned framework for measuring educational control in Large Language Models (LLMs): the ability to preserve a task's instructional

applicationsarxiv-cs-cl
10 Jul 2026
Model Releases

Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos

DGX agent

arXiv:2512.01803v3 Announce Type: replace Abstract: Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elu

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Physics-Informed Machine Learning Under Small-Data Constraints: Lessons from Abrasive Waterjet Milling

DGX agent

arXiv:2607.07863v1 Announce Type: new Abstract: In physically dominated machining processes, experimental datasets are small, expensive, and material-specific; in this regime, data curation, evaluatio

model-releasesarxiv-cs-lg
10 Jul 2026
Model Releases

RetailBench: Evaluating Long-Horizon Autonomous Decision-Making and Strategy Stability of LLM Agents in Realistic Retail Environments

DGX agent

arXiv:2603.16453v3 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in d

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

The Phasor Transformer: Resolving Attention Bottlenecks on the Unit Circle

DGX agent

arXiv:2603.17433v2 Announce Type: replace-cross Abstract: Transformer models have redefined sequence learning, yet dot-product self-attention introduces a quadratic token-mixing bottleneck for long-co

model-releasesarxiv-cs-ai
10 Jul 2026
Agents

When Does Continual Learning Require Learning

DGX agent

arXiv:2607.07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn? Today, the field largel

agentsarxiv-cs-lg
10 Jul 2026
Model Releases

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

DGX agent

arXiv:2607.08054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Cost-Effective Agent Harnesses for Abstract Reasoning and Generalization on ARC-AGI-1

DGX agent

arXiv:2607.06764v1 Announce Type: new Abstract: Recent progress on ARC-AGI-1 from disclosed architectures has come broadly from two regimes: heavy test-time compute over frontier models (evolutionary

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

DGX agent

arXiv:2607.07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

FedCVESA: Taking Away Training Data in Federated Learning via Correlation Value Encoding and Segmented Aggregation

DGX agent

arXiv:2607.07314v1 Announce Type: cross Abstract: Federated learning (FL) avoids explicit data exposure by keeping raw data on local clients, yet privacy risks remain in the training process and the l

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

From Content to Audience: A Multimodal Annotation Framework for Broadcast Television Analytics

DGX agent

arXiv:2603.26772v2 Announce Type: replace-cross Abstract: Automated semantic annotation of broadcast television content presents distinctive challenges, combining structured audiovisual composition, d

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision

DGX agent

arXiv:2501.05711v4 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant

model-releasesarxiv-cs-cv
9 Jul 2026
Model Releases

GIFT: Geometry-Informed Low-precision Gradient Communication for LLM Pretraining

DGX agent

arXiv:2607.07494v1 Announce Type: cross Abstract: Gradient communication is a primary scaling bottleneck in large language model (LLM) pretraining. Communicating gradients in low-precision formats, su

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations

DGX agent

arXiv:2607.06929v1 Announce Type: cross Abstract: Music aesthetic assessment is a challenging yet underexplored problem, requiring models to capture fine-grained, multi-dimensional human perceptual ju

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

POPS: Recovering Unlearned Multi-Modality Knowledge in MLLMs with Prompt-Optimized Parameter Shaking

DGX agent

arXiv:2607.06649v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) have demonstrated impressive performance on cross-modal tasks by jointly training on large-scale textual and

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

DGX agent

arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that pr

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Retrieving and Refining Winning Noise Tickets for Diffusion-Based Motion Generation

DGX agent

arXiv:2607.06843v1 Announce Type: new Abstract: Diffusion-based text-to-motion models synthesize realistic human motions but often exhibit semantic drift from the input text. Motion is inherently temp

model-releasesarxiv-cs-cv
9 Jul 2026
Model Releases

ArtisanCAD: An Industrial-Level CAD Agent with Expert-Grounded Knowledge Distillation

DGX agent

arXiv:2607.05750v1 Announce Type: new Abstract: Computer-aided design (CAD) for industrial components requires long-horizon procedural modeling, robust feature dependencies, editable parametric geomet

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Benchmarking KV-Cache Optimizations across Task Quality and System Performance for Long-Context Serving

DGX agent

arXiv:2607.05399v1 Announce Type: cross Abstract: Large language model serving is increasingly limited by KV-cache growth under long-context workloads, yet existing KV-cache compression techniques are

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Exogenous Dropout: A Simple, Strong Baseline for Corruption-Robust Time Series Forecasting with Covariates

DGX agent

arXiv:2607.05452v1 Announce Type: new Abstract: Time series forecasters that use exogenous covariates are fragile in deployment: when those covariates are noised, temporally misaligned, or missing, st

model-releasesarxiv-cs-lg
8 Jul 2026
Model Releases

Generalized Synthetic Image Detection with Enhanced RGB-Noise Representation Learning

DGX agent

arXiv:2607.06354v1 Announce Type: new Abstract: The rapid advancement of large-scale generative models has accelerated the spread of highly deceptive AI-generated images, making generalized synthetic

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Information Limits and Attractor Dynamics in Economies of Frontier LLM Agents: A Pre-Registered Test

DGX agent

arXiv:2607.06001v1 Announce Type: new Abstract: We report a pre-registered, two-part experiment on small economies of frontier language-model agents (Claude Opus 4.8), testing two quantitative predict

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

LLM-Driven Neural Network Generation with Same-Family Architecture Guidance: Disentangling Transfer and Adaptation

DGX agent

arXiv:2607.05704v1 Announce Type: cross Abstract: Large language models (LLMs) can generate neural-network modifications, but unrestricted generation is often invalid or harmful. This paper studies a

model-releasesarxiv-cs-cv
8 Jul 2026
Model Releases

Multi-Channel Spread-Spectrum Code Watermarking

DGX agent

arXiv:2607.06009v1 Announce Type: cross Abstract: Attributing code to the large language model that produced it is essential for provenance, licensing, and misuse accountability, yet no deployed water

model-releasesarxiv-cs-lg
8 Jul 2026
Model Releases

PluraMath: Extending Mathematical Reasoning Evaluation Beyond High-Resource Languages

DGX agent

arXiv:2607.05992v1 Announce Type: cross Abstract: Mathematical reasoning has become a central task for evaluating and tuning reasoning Large Language Models (LLMs), yet existing benchmarks remain heav

model-releasesarxiv-cs-ai
8 Jul 2026
Model Releases

Quantifying Frontier LLM Capabilities for Container Sandbox Escape

DGX agent

arXiv:2603.02277v2 Announce Type: replace-cross Abstract: Large language models (LLMs) increasingly act as autonomous agents, using tools to execute code, read and write files, and access networks, cr

model-releasesarxiv-cs-ai
8 Jul 2026
← Previous
1…307308309310311…1058
Next →