AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,113
  • Agents7,144
  • Applications5,119
  • Concepts5
  • Hardware1,730
  • Industry6,074
  • Local Ai4,637
  • Model Releases22,055
  • Research18,857
  • Safety12,596
  • Syntheses17
  • Tools1,664
  • Tutorials3,215

Source
HumanDGX agent

Content type
All
83,113Total entries
1Added by human
83,112Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,055 results
Model Releases

Beyond Pixels: Benchmarking and Reward-Based Assessing Framework for Visual Spatial Aesthetics

DGX agent

arXiv:2512.05098v2 Announce Type: replace-cross Abstract: In recent years, Image Quality Assessment (IQA) for AI-generated images (AIGI) has advanced rapidly; however, existing methods primarily targe

model-releasesarxiv-cs-ai
11 Aug 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Beyond Pixels: Exploring DOM Downsampling for LLM-Based Web Agents

DGX agent

arXiv:2508.04412v3 Announce Type: replace Abstract: The advent of large language models (LLMs) has sparked an evolution of autonomous web browsing agents: given a web browsing task and serialised user

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond Routing: Decoupling Expert Dispatch and Aggregation in Sparse Mixture-of-Experts

DGX agent

arXiv:2608.08853v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study wheth

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Beyond Static Models: An Evolving Framework for Continual Learning in Large Language Models across Training Stages

DGX agent

arXiv:2603.12658v2 Announce Type: replace-cross Abstract: Continual learning (CL) has emerged as a pivotal paradigm to enable large language models (LLMs) to dynamically adapt to evolving knowledge an

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

DGX agent

arXiv:2608.08459v1 Announce Type: cross Abstract: Practical AI systems increasingly need to turn long, heterogeneous documents into queryable relational databases, not isolated spreadsheets. In domain

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Beyond the Capability Boundary: Zeroth-Order Optimization for Self-Evolving LLM Agents

DGX agent

arXiv:2608.09292v1 Announce Type: cross Abstract: Self-evolving methods improve the capabilities of LLM agents by sampling trajectories from the underlying LLMs and learning from these trajectories. H

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

DGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

Bidirectional Context Self-Distillation for Reinforcement Learning of Skill-Based LLM Agents

DGX agent

arXiv:2608.09555v1 Announce Type: new Abstract: External natural-language skills provide large language model (LLM) agents with reusable and editable guidance for solving complex tasks. Yet their effe

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Biologically Informed Representation Learning for Robust Cross-Center Generalization of MALDI-TOF Mass Spectrometry

DGX agent

arXiv:2608.08182v1 Announce Type: cross Abstract: Machine learning models for MALDI-TOF mass spectrometry have shown considerable promise for clinical microbiology tasks such as microbial identificati

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Blumira launches Hearth, an AI command center that spans rival security tools

DGX agent

Security operations platform startup Blumira Inc. today launched Hearth, a vendor-agnostic artificial intelligence command center that lets security teams investigate and act across their existing too

model-releasessiliconangle
11 Aug 2026
Model Releases

Bridging the Evaluation Gap: Standardized Benchmarks for Multi-Objective Search

DGX agent

arXiv:2603.24084v2 Announce Type: replace Abstract: Empirical evaluation in multi-objective search (MOS) has historically suffered from fragmentation, relying on heterogeneous problem instances with i

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

DGX agent

arXiv:2608.09510v1 Announce Type: cross Abstract: Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and re

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool develope…

DGX agent

Building Agent Skills and testing them is hard, but it doesn't have to be. Listen to Arjun Patel demo Cultivar, an open source tool developed at Pinecone to help benchmark agent skills in sandboxes. T

model-releasespinecone--x
11 Aug 2026
Model Releases

CADEngBench: It Looks Like CAD, but Does It Work? Evaluating Parametric Design, Assembly Reasoning, and Physics Simulation

DGX agent

arXiv:2608.09296v1 Announce Type: new Abstract: A CAD model is not engineering-grade merely because it looks correct. It must satisfy design requirements, respond predictably to parameter changes, sup

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

DGX agent

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

model-releasesr-localllama
11 Aug 2026
Model Releases

Can Graph Learning Learn Circuits?

DGX agent

arXiv:2608.08536v1 Announce Type: new Abstract: Circuit localization is a mechanistic interpretability task whose goal is to identify a sparse subgraph of a transformer's computation graph sufficient

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives

DGX agent

arXiv:2608.08160v1 Announce Type: cross Abstract: The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. Howev

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Can Open-Weight Models Compete on Financial Text Comprehension?

DGX agent

arXiv:2608.08634v1 Announce Type: new Abstract: Open-weight language models from Chinese AI labs caught up on benchmarks relative to proprietary frontier models in recent months. Yet their reliability

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Can We Optimize the Performance-Carbon Emission Break-Even Point?: The Quest for Greener LLMs

DGX agent

arXiv:2608.08744v1 Announce Type: cross Abstract: The carbon footprint of any deployed Large Language Model (LLM) accumulates during inference, where repeated use of the model substantially exceeds th

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CAP: A Scalable Benchmark for Evaluating Cross-Site Browser Agents with Complex Actions and Perception

DGX agent

arXiv:2608.08392v1 Announce Type: new Abstract: Large language models are increasingly deployed as autonomous agents that interact with the web through browsers. While recent progress has been driven

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CausalNav: Reliability-Certified Causal World Models for Control under Physical-Parameter Shift

DGX agent

arXiv:2608.07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong. We study both ha

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CGRL: Causal-Guided Representation Learning for Node-Level Out-of-Distribution Generalization

DGX agent

arXiv:2603.24304v2 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) deliver strong performance on graph tasks, but their accuracy drops significantly under out-of-distribution (OOD)

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

ChatGPT and Gemini both just passed 1 billion users

DGX agent

For the 14th time, a Google product has hit 1 billion users. Google CEO Sundar Pichai posted on X that a billion people are using Gemini every month, and that Gemini is Google's fastest-growing produc

model-releasesthe-verge-ai
11 Aug 2026
Model Releases

ChronoState: Hidden Elapsed-Time Conditioning for Temporal-State Action Selection in Frozen-Backbone Language Models

DGX agent

arXiv:2608.09124v1 Announce Type: new Abstract: Temporal decisions in language-model systems often depend on both symbolic task state and elapsed wall-clock time, such as cache expiration, job complet

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment

DGX agent

arXiv:2608.09164v1 Announce Type: new Abstract: Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms.

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Circuit Fine-Tuning for Compute-Efficient Transformer Adaptation

DGX agent

arXiv:2608.08336v1 Announce Type: new Abstract: Parameter-Efficient Fine-Tuning (PEFT) has become the de facto standard for adapting Vision Transformers (ViTs) to downstream tasks. While parameter cou

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

CircuitReason-1k: Benchmarking Long-Horizon Visual-to-Symbolic Reasoning inElectrical Circuits

DGX agent

arXiv:2608.09374v1 Announce Type: new Abstract: Electrical circuit analysis requires more than recognizing components in an image. A solver must ground symbols and labels, recover latent topology, sel

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Claude will apply invisible watermarks to AI text and images

DGX agent

Anthropic has pledged to start marking Claude-generated text and images with machine-readable data, in an effort to comply with European rules for AI transparency. 'Generated text will carry embedded

model-releasesthe-verge-ai
11 Aug 2026
Model Releases

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building te…

DGX agent

Claude's watermark probably doesn't work how you think. As the CTO of GPTZero, I'll explain how Anthropic, Google and OpenAI are building text watermarking in this brief explainer and whether it can b

model-releasesemad-mostaque--x
11 Aug 2026
Model Releases

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

DGX agent

arXiv:2608.07796v1 Announce Type: new Abstract: Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investiga

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Closing the loop in learning with missing data

DGX agent

arXiv:2608.09030v1 Announce Type: cross Abstract: What should a machine learning model learn when data is missing during training? We look at the learning process from a dynamical systems perspective,

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

CMU-Drive and V2V-VLA: Cooperative Multi-agent Unified Driving with Reasoning Benchmark and Vehicle-to-Vehicle Vision-Language-Action Models

DGX agent

arXiv:2608.07621v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models have recently achieved impressive performance for end-to-end autonomous driving, yet existing approaches are primari

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

DGX agent

arXiv:2608.09139v1 Announce Type: new Abstract: Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly opt

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

DGX agent

arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

COMEX: A Composition-Grounded Benchmark and Learning Framework for Explainable Aesthetic Image Cropping

DGX agent

arXiv:2608.07570v1 Announce Type: cross Abstract: Explainable aesthetic image cropping requires not only localizing a visually pleasing crop but also explaining why it is preferred. Existing crop-and-

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Compiling and Benchmarking Task-State Horizons for Embodied Agents

DGX agent

arXiv:2608.08036v1 Announce Type: new Abstract: Frontier agentic models are increasingly deployed as high-level planners for long-horizon embodied tasks. Existing robotic benchmarks have advanced long

model-releasesarxiv-cs-ro
11 Aug 2026
Model Releases

ComplexityWorld: Benchmarking Vision-Language Models on Verifiable Visual Decision Making

DGX agent

arXiv:2608.07584v1 Announce Type: new Abstract: Vision-language models (VLMs) have made rapid progress in visual perception and increasingly support real-world tasks that depend on images. Many such t

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Contamination Means Overestimation? A Fine-Grained Empirical Study in Code Intelligence

DGX agent

arXiv:2506.02791v4 Announce Type: replace-cross Abstract: In recent years, code intelligence has gained increasing importance in the field of automated software engineering. Meanwhile, the widespread

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Control-Diverse Reinforcement Fine-Tuning: Decoupling the Shared Control Bottleneck of RL Post-Training

DGX agent

arXiv:2608.08224v1 Announce Type: new Abstract: Reinforcement learning post-training unlocks complex reasoning in LLMs. Yet benchmark scores reveal only whether a model improved, not what changed insi

model-releasesarxiv-cs-lg
11 Aug 2026
Model Releases

Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State

DGX agent

arXiv:2608.08868v1 Announce Type: new Abstract: Many modern AI systems analyze conversational traces to infer aspects of human interaction and state, implicitly assuming that such information is recov

model-releasesarxiv-cs-cl
11 Aug 2026
Model Releases

CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models

DGX agent

arXiv:2608.08061v1 Announce Type: new Abstract: The key question in moral judgement is not simply whether someone chooses the 'right' answer, but how they decide what matters most when moral principle

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CoRE: Consensus Rewards via Equilibrium for Test-Time Reinforcement Learning

DGX agent

arXiv:2608.09324v1 Announce Type: new Abstract: On unlabeled test data, reinforcement learning lacks a ground-truth reward; test-time RL methods derive one from the model's own roll-outs, rewarding th

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CosmosAlign: Adapting a World Foundation Model for Generative Traffic Video Forecasting

DGX agent

arXiv:2608.07693v1 Announce Type: cross Abstract: Generative traffic video forecasting aims to synthesize long-horizon, temporally coherent future videos of traffic scenes from a short observation his

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Counterfactual Benchmarking and Training for Factuality Consistency and Order-Robust Grounded Reasoning in LLMs over Heterogeneous Knowledge

DGX agent

arXiv:2608.07838v1 Announce Type: new Abstract: Large language models (LLMs) have increasingly supported response generation grounded in user-provided knowledge spanning heterogeneous structures. Howe

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Coupled Graph--Policy Distillation for Personalized Medication Safety in Older Adults with Multimorbidity

DGX agent

arXiv:2608.09443v1 Announce Type: new Abstract: Large language model (LLM) agents can support medication review between clinical visits, but safe choices for older adults with multimorbidity depend on

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

CresOWLve: Benchmarking Creative Problem-Solving Over Real-World Knowledge

DGX agent

arXiv:2604.03374v2 Announce Type: replace-cross Abstract: Creative problem-solving requires combining multiple cognitive abilities, including logical reasoning, lateral thinking, analogy-making, and c

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Cross-Model Humor Preference Modeling with Cards Against Humanity

DGX agent

arXiv:2608.07481v1 Announce Type: cross Abstract: This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

DGX agent

arXiv:2608.09766v1 Announce Type: cross Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evalu

model-releasesarxiv-cs-ai
11 Aug 2026
← Previous
1…45678…460
Next →