AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,793 results
Model Releases

The Multilingual Quantization Tax: Structural Collapse and Typological Fragility in Edge SLMs

DGX agent

arXiv:2608.09941v1 Announce Type: new Abstract: While 4-bit weight quantization is critical for deploying Small Language Models (SLMs) on edge devices, evaluations of the resulting performance degrada

model-releasesarxiv-cs-cl
12 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Try Grok 4.6 on tough real-world tasks!

DGX agent

Try Grok 4.6 on tough real-world tasks! imo GDPVal is probably the most important benchmark, it measures the performance of models on real world tasks Big leap in performance here to top it at a great

model-releaseselon-musk--x
12 Aug 2026
Model Releases

Uncertainty-Aware Ensemble Deep Randomized Neural Networks for Classification

DGX agent

arXiv:2608.10007v1 Announce Type: cross Abstract: The current state-of-the-art (SOTA) deep randomized neural networks, such as deep Random Vector Functional Link (dRVFL) and ensemble deep RVFL (edRVFL

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

V-FiLLM: Verified Financial LLM Reasoning Benchmark

DGX agent

arXiv:2608.11047v1 Announce Type: new Abstract: While existing benchmarks have made substantial progress in evaluating LLMs across STEM domains, financial reasoning over structured data remains compar

model-releasesarxiv-cs-ai
12 Aug 2026
Agents

VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection

DGX agent

arXiv:2511.19436v2 Announce Type: replace-cross Abstract: Existing Video Detailed Captioning (VDC) methods predominantly rely on costly human annotations or distillation from powerful proprietary mode

agentsarxiv-cs-ai
12 Aug 2026
Safety

What DINO saw: ALiBi positional encoding reduces positional bias in Vision Transformers

DGX agent

arXiv:2603.16840v2 Announce Type: replace Abstract: Vision transformers (ViTs) - especially feature foundation models like DINOv2 - learn rich representations useful for many downstream tasks. However

safetyarxiv-cs-cv
12 Aug 2026
Model Releases

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

DGX agent

arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of

model-releasesarxiv-cs-ai
12 Aug 2026
Model Releases

A Unified Issue Resolution Benchmark for Requirement Clarification, Planning, and Code Generation for Coding Agents

DGX agent

arXiv:2608.09072v1 Announce Type: cross Abstract: Large language model-powered coding agents are increasingly used to modify existing code repositories, for example, by adding features or fixing bugs.

model-releasesarxiv-cs-ai
11 Aug 2026
Industry

Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock

DGX agent

Daybreak Red and Daybreak Blue from OpenAI, specialized cyber defense models from OpenAI, are now available on Amazon Bedrock to eligible customers. Both models run with zero-operator access enforced

industryaws-ml-blog
11 Aug 2026
Model Releases

ACEvo: Adversarial Co-Evolution of Problem Distributions and Solvers for Combinatorial Optimization

DGX agent

arXiv:2506.02594v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to synthesize heuristic programs, yet most existing pipelines optimize solvers against fixed benc

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

ActBench: Self-Evolving Benchmark of Behavioral Safety in Cowork Agents

DGX agent

arXiv:2608.09476v1 Announce Type: cross Abstract: Cowork agents may complete benign tasks while disclosing protected data, manipulating unauthorized state, invocate unauthorized API. We define behavio

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

ADOPD: Reference-Privileged On-Policy Distillation for MLLM-Based Industrial Anomaly Detection

DGX agent

arXiv:2608.09789v1 Announce Type: new Abstract: Industrial anomaly detection (IAD) requires identifying fine-grained deviations from normal visual patterns. Multimodal large language models (MLLMs) ca

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

ASPaeroFlow: Decomposition Heuristics for Joint Air Traffic Flow & Capacity Management

DGX agent

arXiv:2608.09315v1 Announce Type: new Abstract: While mathematical models act as vital decision support systems for operational Air Traffic Flow and Capacity Management (ATFCM), existing approaches is

model-releasesarxiv-cs-ai
11 Aug 2026
Applications

Backward Compatibility in Tree-Based Explanations and Enhanced CART Algorithm

DGX agent

arXiv:2608.08674v1 Announce Type: new Abstract: In the operation of machine learning models, model update is a fundamental process that requires careful consideration of its impact on downstream decis

applicationsarxiv-cs-lg
11 Aug 2026
Research

Beyond Global Editing: Per-Instance Disentangled Subspaces for Training-Free Hallucination Mitigation in LVLMs

DGX agent

arXiv:2608.09344v1 Announce Type: new Abstract: Recent advances in large vision-language models (LVLMs) have enabled powerful multimodal reasoning by integrating visual encoders with large language mo

researcharxiv-cs-cv
11 Aug 2026
Model Releases

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

DGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

DGX agent

arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effe

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

DGX agent

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup

model-releasesarxiv-cs-ai
11 Aug 2026
Research

Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents

DGX agent

arXiv:2608.09485v1 Announce Type: new Abstract: Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission,

researcharxiv-cs-ai
11 Aug 2026
Model Releases

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

DGX agent

arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, sel

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

DGX agent

arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investiga

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Closing the loop in learning with missing data

DGX agent

arXiv:2608.09030v1 Announce Type: cross Abstract: What should a machine learning model learn when data is missing during training? We look at the learning process from a dynamical systems perspective,

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence

DGX agent

arXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

DGX agent

arXiv:2608.09324v1 Announce Type: new Abstract: On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding th

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

DGX agent

arXiv:2608.09766v1 Announce Type: cross Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evalu

model-releasesarxiv-cs-ai
11 Aug 2026
Local Ai

Decided Upstream, Written Late: Locating and Pricing the Cross-Lingual Refusal Circuit of a Multilingual MoE

DGX agent

arXiv:2608.08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in

local-aiarxiv-cs-ai
11 Aug 2026
Hardware

Design Space of Self--Consistent Electrostatic Machine Learning Interatomic Potentials

DGX agent

arXiv:2603.14700v2 Announce Type: replace-cross Abstract: Machine learning interatomic potentials (MLIPs) have become widely used tools in atomistic simulations. For much of the history of this field,

hardwarearxiv-cs-lg
11 Aug 2026
Model Releases

DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference

DGX agent

arXiv:2608.08878v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequen

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

DocAtlas: Long-Document Understanding as Mutable-State Interaction

DGX agent

arXiv:2608.07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-a

model-releasesarxiv-cs-ai
11 Aug 2026
Research

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees

DGX agent

arXiv:2608.08002v1 Announce Type: new Abstract: Language-model judges and reward models enable scalable supervision, but finite optimization can exploit evaluator errors rather than improve response q

researcharxiv-cs-lg
11 Aug 2026
Model Releases

FaLCon: Facet-Anchored Retrieval with Late Consensus for Sim2Real Text-Based Person Anomaly Search

DGX agent

arXiv:2608.09474v1 Announce Type: new Abstract: Text-based person anomaly search requires retrieving real-world pedestrian images from detailed natural-language descriptions using models trained prima

model-releasesarxiv-cs-cv
11 Aug 2026
Safety

Financial Numerical Prediction and Allocation as Token Generation

DGX agent

arXiv:2608.09880v1 Announce Type: new Abstract: Financial prediction typically relies on task-specific regression, ranking, or policy heads, separating the language model from the numerical object ult

safetyarxiv-cs-cv
11 Aug 2026
Model Releases

ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration

DGX agent

arXiv:2608.08605v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common ba

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

From Manuals to Maintenance: Fine-Tuning MedGemma for Multi-Modal Imaging System Support in Low-Resource Settings

DGX agent

arXiv:2608.08896v1 Announce Type: new Abstract: Imaging device downtime is a major barrier to healthcare delivery in low- and middle-income countries (LMICs), often driven by limited access to special

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

GENCO - A Unified Neural Solver Embedded in a Development Framework for Steady-State Grid Analysis

DGX agent

arXiv:2608.09921v1 Announce Type: new Abstract: Foundation models are transforming business workflows and boosting productivity, yet they remain largely absent from engineering domains such as power s

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Goedel-Code-Prover: Hierarchical Proof Search for Open State-of-the-Art Code Verification

DGX agent

arXiv:2603.19329v3 Announce Type: replace-cross Abstract: Large language models (LLMs) can generate plausible code but offer limited guarantees of correctness. Formally verifying that implementations

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

DGX agent

arXiv:2608.08466v1 Announce Type: new Abstract: Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harne

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

High-Layer Attention Pruning with Rescaling

DGX agent

arXiv:2507.01900v3 Announce Type: replace Abstract: Pruning is a highly effective approach for compressing large language models (LLMs), significantly reducing inference latency. However, conventional

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation

DGX agent

arXiv:2608.07498v1 Announce Type: cross Abstract: Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deploym

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?

DGX agent

arXiv:2503.19990v4 Announce Type: replace Abstract: Many real-world applications of spatial intelligence, such as robotic control, autonomous driving, and automated assembly, require spatial reasoning

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

DGX agent

arXiv:2607.15509v2 Announce Type: replace-cross Abstract: We present a fully automated closed-loop AutoML framework that uses GPT-5, GPT-4o, and Claude Sonnet 4 as autonomous neural architecture desig

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks

DGX agent

arXiv:2507.03162v2 Announce Type: replace-cross Abstract: The rapid advancement of Large Language Models (LLMs) has transformed various domains, particularly computer science (CS) education. These mod

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Measuring the Tokenization Premium: A Cost Audit for Underserved Language Communities

DGX agent

arXiv:2608.09046v1 Announce Type: new Abstract: Large language models are increasingly deployed as general-purpose educational and technical assistance systems, but their underlying infrastructure doe

model-releasesarxiv-cs-cl
11 Aug 2026
Research

MixFormer: Linear Transformer with Mixture of Memory Experts

DGX agent

arXiv:2608.09468v1 Announce Type: cross Abstract: State Space Models (SSMs), as a mainstream research direction of linear Transformers, aim to achieve higher efficiency than standard Transformers in l

researcharxiv-cs-ai
11 Aug 2026
Model Releases

MRBench: A Comprehensive Benchmark for Human Motion-Text Retrieval

DGX agent

arXiv:2608.07993v1 Announce Type: new Abstract: Human motion-text retrieval provides a rigorous means of assessing cross-modal alignment. Prevailing benchmarks are dominated by homogeneous indoor moti

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

NeuroAda: Activating Each Neuron's Potential for Parameter-Efficient Fine-Tuning

DGX agent

arXiv:2510.18940v2 Announce Type: replace-cross Abstract: Existing parameter-efficient fine-tuning (PEFT) methods primarily fall into two categories: addition-based and selective in-situ adaptation. T

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Neurosymbolic Discovery of Algebraic Graph Constructions

DGX agent

arXiv:2608.08118v1 Announce Type: new Abstract: There are several methods for searching for graphs with prescribed properties, such as SAT solvers and specialized generators. These methods return the

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

OBLIVION: Workflow-Level Operational Skill Unlearning for Deployed Agents

DGX agent

arXiv:2608.08264v1 Announce Type: new Abstract: Large language model agents are becoming operational interfaces to files, memories, registries, and external tools. This deployment shift creates a new

model-releasesarxiv-cs-ai
11 Aug 2026
← Previous
1…435436437438439…1371
Next →