AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,793 results
Model Releases

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

DGX agent

arXiv:2608.09873v1 Announce Type: cross Abstract: We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It co

model-releasesarxiv-cs-ai
11 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

SciTaRC: A Plan-Annotated Scientific Tabular QA Benchmark for Language Reasoning and Complex Computation

DGX agent

arXiv:2603.08910v2 Announce Type: replace Abstract: We introduce SciTaRC, an expert-authored benchmark for question answering over scientific tables that targets composite, multi-step reasoning. To en

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Shape Mutating Expert Compression:LorExperts and BTExperts

DGX agent

arXiv:2608.07814v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) language models deliver high capacity at low per-token compute, but deploying them cheaply requires compressing their many ex

model-releasesarxiv-cs-ai
11 Aug 2026
Tutorials

SignLlama: Enhancing Gloss-free Sign Language Translation by Prioritizing Visual Features for LLMs

DGX agent

arXiv:2608.09006v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved remarkable success across a wide range of tasks. However, fine-tuning LLMs for Gloss-Free Sign Language Tra

tutorialsarxiv-cs-ai
11 Aug 2026
Model Releases

Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

DGX agent

arXiv:2608.09128v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and imp

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

DGX agent

arXiv:2608.07862v1 Announce Type: new Abstract: Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguist

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators

DGX agent

arXiv:2608.07641v1 Announce Type: new Abstract: The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code Refactoring

DGX agent

arXiv:2608.09802v1 Announce Type: new Abstract: As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluati

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

TongGuOCR: A Layout-Aware and Token-Augmented OCR Framework for Chinese Historical Documents

DGX agent

arXiv:2608.07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-te

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

TREAT: Evaluating Access to Formal Knowledge across Equivalent Mathematical Representations

DGX agent

arXiv:2608.07540v1 Announce Type: new Abstract: AI systems increasingly operate between flexible input representations and formal objects used by downstream tools. A key challenge is recognizing when

model-releasesarxiv-cs-ai
11 Aug 2026
Safety

Understanding Reasoning from Pretraining to Post-Training

DGX agent

arXiv:2607.16097v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is l

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

Verication-driven closed-loop multi-agent large language modelframework for code-compliant structural design

DGX agent

arXiv:2608.07978v1 Announce Type: cross Abstract: Multi-agent large language model(LLM)systems are applied to structural design,yet most use one-shot generation and cannot verify their output,leaving

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning Systems

DGX agent

arXiv:2401.04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of in

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks

DGX agent

arXiv:2506.01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent

model-releasesarxiv-cs-ai
11 Aug 2026
Local Ai

When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs

DGX agent

arXiv:2608.08159v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lin

local-aiarxiv-cs-ai
11 Aug 2026
Model Releases

Zero-Shot Traffic Accident Detection via a Coarse-to-Fine VLM-Tracking Pipeline

DGX agent

arXiv:2608.08867v1 Announce Type: new Abstract: Traffic surveillance cameras capture accidents continuously, yet converting raw CCTV footage into structured event records that pinpoint when, where, an

model-releasesarxiv-cs-cv
11 Aug 2026
Research

Adversarial Causal Intervention Falsification

DGX agent

arXiv:2608.06427v1 Announce Type: new Abstract: Generative models can reproduce an observational distribution while encoding an incorrect causal structure. We study a sequential game in which a struct

researcharxiv-cs-lg
10 Aug 2026
Model Releases

Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications

DGX agent

arXiv:2608.07250v1 Announce Type: cross Abstract: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will redu

model-releasesarxiv-cs-ai
10 Aug 2026
Safety

Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques

DGX agent

arXiv:2608.06609v1 Announce Type: new Abstract: Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testin

safetyarxiv-cs-ai
10 Aug 2026
Local Ai

Boundary Density Likelihood for Direct Event-Time Supervision

DGX agent

arXiv:2408.12792v2 Announce Type: replace Abstract: Event detection turns long recordings into a sparse set of ranked timestamps. Yet many sequence models are trained for samplewise segmentation and o

local-aiarxiv-cs-ai
10 Aug 2026
Research

Counterfactual Simulation Training for Chain-of-Thought Faithfulness

DGX agent

arXiv:2602.20710v2 Announce Type: replace Abstract: Inspecting Chain-of-Thought reasoning is among the most common means of understanding why an LLM produced its output. But well-known problems with C

researcharxiv-cs-ai
10 Aug 2026
Model Releases

ED-CSP: Crystal Structure Prediction from Electron Diffraction

DGX agent

arXiv:2608.06448v1 Announce Type: cross Abstract: Recovering a periodic 3D crystal structure from sparse, unindexed electron diffraction (ED) observations is a challenging generative inverse problem.

model-releasesarxiv-cs-ai
10 Aug 2026
Tutorials

Embedded Variational Neural Stochastic Differential Equations for Learning Heterogeneous Dynamics

DGX agent

arXiv:2604.00669v2 Announce Type: replace Abstract: This study examines the challenges of modeling complex and noisy data related to socioeconomic factors over time, with a focus on data from various

tutorialsarxiv-cs-lg
10 Aug 2026
Agents

Explicit, Not Longer: What Makes Epistemic Stance Survive Memory Compression

DGX agent

arXiv:2608.06953v1 Announce Type: cross Abstract: Agent memory systems compress what they store, and compression is built to drop qualifiers, so a claim's epistemic standing tends not to survive being

agentsarxiv-cs-ai
10 Aug 2026
Research

FedLBW: A Loss-Based Weighting Strategy for Federated Learning on Non-IID Data in Wireless Networks

DGX agent

arXiv:2608.07007v1 Announce Type: new Abstract: Federated Learning (FL) enables collaborative machine learning (ML) across distributed clients while preserving privacy. However, efficient model conver

researcharxiv-cs-ai
10 Aug 2026
Model Releases

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

DGX agent

arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t

model-releasesarxiv-cs-ai
10 Aug 2026
Safety

From Cheap Fakes to Pure Synthesis: Addressing the New Era of T2V Fake News Videos

DGX agent

arXiv:2608.06732v1 Announce Type: new Abstract: Recent text-to-video (T2V) generation models enable fake news videos to be synthesized from scratch, shifting the threat beyond cheap fakes assembled fr

safetyarxiv-cs-ai
10 Aug 2026
Model Releases

GeoBenchLLM: A Comprehensive Benchmark for Evaluating LLMs on Geo-Related Tasks

DGX agent

arXiv:2608.07411v1 Announce Type: new Abstract: In the context of geodata, existing Large Language Models have often been studied in a homogeneous setting, which has considerably limited insights into

model-releasesarxiv-cs-ai
10 Aug 2026
Research

GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base

DGX agent

arXiv:2608.06992v1 Announce Type: cross Abstract: We present a web demo for exploring a large-scale disambiguated knowledge base (KB) materialized from a large language model (LLM). GPTKB 2.0 contains

researcharxiv-cs-ai
10 Aug 2026
Local Ai

HazeSpikeMamba: Coupling Spiking-Inspired and State-Space Features for Self-Supervised Real-World Dehazing

DGX agent

arXiv:2608.06886v1 Announce Type: new Abstract: Dehazing networks are commonly trained on synthetic hazy-clear pairs, but their performance often drops on real photographs. Synthetic haze generated us

local-aiarxiv-cs-cv
10 Aug 2026
Safety

How WPP operationalizes platform and data engineering for AI marketing

DGX agent

Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and op

safetygoogle-cloud-ai
10 Aug 2026
Research

International Transfer of Stochastic Cortical Self-Reconstruction

DGX agent

arXiv:2608.07092v1 Announce Type: cross Abstract: Stochastic cortical self-reconstruction (SCSR) enables personalized mapping of gray matter atrophy, a hallmark of neurodegenerative disorders such as

researcharxiv-cs-ai
10 Aug 2026
Model Releases

LitTraceQA: A Benchmark for Multi-Stage Grounding and Verification in Scientific Question Answering

DGX agent

arXiv:2608.07370v1 Announce Type: new Abstract: Scientific literature is increasingly used as a knowledge source for language models, retrieval-augmented generation systems, and research assistants, b

model-releasesarxiv-cs-cl
10 Aug 2026
Safety

LMM Modality Transfer: A Pre-requisite for Autonomous GIS Agents

DGX agent

arXiv:2608.06948v1 Announce Type: new Abstract: AI models are becoming increasingly adept at understanding and processing spatial information, thereby facilitating agentic problem-solving in spatial t

safetyarxiv-cs-ai
10 Aug 2026
Model Releases

LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer's Disease Screening

DGX agent

arXiv:2608.07378v1 Announce Type: cross Abstract: Early diagnosis of Alzheimer's disease (AD) is critical for enabling timely interventions that may slow disease progression and improve patient outcom

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

MAC: A Conversion Rate Prediction Benchmark Featuring Labels Under Multiple Attribution Mechanisms

DGX agent

arXiv:2603.02184v2 Announce Type: replace-cross Abstract: Multi-attribution learning (MAL), which enhances model performance by learning from conversion labels yielded by multiple attribution mechanis

model-releasesarxiv-cs-ai
10 Aug 2026
Model Releases

Motif-Technologies/Motif-3 official realese

DGX agent

Motif-Technologies is one of the tech company participated South Korea's AI Foundation Model project.(독파모) Upstage(Solar Series), LG AI Research(EXAONE Series), and SKT(A.X Series) are the competitors

model-releasesr-localllama
10 Aug 2026
Model Releases

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, bec…

DGX agent

Oh no, we aren’t going to go back to this sort of prompting again, are we? I would love Anthropic to test if it actually works robustly, because our experiments (with slightly older models) found it d

model-releasesethan-mollick--x
10 Aug 2026
Research

oldsymbol{lambda}-Orthogonality Regularization for Compatible Representation Learning

DGX agent

arXiv:2509.16664v2 Announce Type: cross Abstract: Retrieval systems rely on representations learned by increasingly powerful models. However, due to the high training cost and inconsistencies in learn

researcharxiv-cs-cv
10 Aug 2026
Model Releases

Please Share Your Experience About Muse Glimmer

DGX agent

I have a classic test for local LLM's. I asked for 8 ball pool game with only one HTML file and Muse Glimmer spend 21k Token(I m using full context so 128k) and only created a 220 lines of HTML and sa

model-releasesr-localllama
10 Aug 2026
Research

Progressive Content Refinement with Decaying Reward Joint LinUCB

DGX agent

arXiv:2608.06750v1 Announce Type: cross Abstract: Iterative refinement has significantly enhanced Large Language Model (LLM) performance; however, existing methods ranging from feedback-based Self-Ref

researcharxiv-cs-ai
10 Aug 2026
Model Releases

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

DGX agent

arXiv:2608.07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters,

model-releasesarxiv-cs-ai
10 Aug 2026
Research

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination

DGX agent

arXiv:2608.07302v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) often suffer from object hallucination, generating objects that are absent from the image. Prior work largely att

researcharxiv-cs-ai
10 Aug 2026
Tutorials

Stochastic Autoregressive Learning

DGX agent

arXiv:2608.07224v1 Announce Type: new Abstract: Motivated by LLMs, which generate outputs by iteratively sampling from next-token distributions, we introduce a PAC-learning model for binary stochastic

tutorialsarxiv-cs-lg
10 Aug 2026
Applications

Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools

DGX agent

arXiv:2608.07446v1 Announce Type: cross Abstract: Rapid adoption of large language models (LLMs) in enterprise settings has introduced operational, security, and governance risks. As generative AI app

applicationsarxiv-cs-ai
10 Aug 2026
Model Releases

Tested Muse Glimmer locally on coding with OpenCode & agentic work

DGX agent

Ran the model with quants (Q4) by Unsloth with latest (build from master) llama.cpp server. It takes ~20GB ram running on M5 Pro with 48GB at about 17t/s. Didn't do any reasoning loops/overthinking. O

model-releasesr-localllama
10 Aug 2026
Safety

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

DGX agent

arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour

safetyarxiv-cs-cl
10 Aug 2026
Model Releases

TRACE: A Multi-Layer Benchmark for Human AI Controller Coordination Under Drift and Failure

DGX agent

arXiv:2608.06657v1 Announce Type: new Abstract: Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustw

model-releasesarxiv-cs-ai
10 Aug 2026
← Previous
1…497498499500501…1371
Next →