AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
90,914Total entries
1Added by human
90,913Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,675 results
Model Releases

PhysMirror: Physics-Aware Mirror Object Generation

DGX agent

arXiv:2607.03470v1 Announce Type: new Abstract: Synthesizing physically accurate mirror reflections remains a fundamental challenge for modern text-to-image diffusion models, which are increasingly cr

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

PosterHarness: Turning Scientific Poster Generation into an Auditable Instruction-Following Benchmark

DGX agent

arXiv:2607.03006v1 Announce Type: cross Abstract: Text-rich image models can now design poster-scale layouts, but we lack ways to measure whether they honor scientific communication contracts: legible

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

Probing Identity-Specific Motion Signatures: A Controlled Diagnostic Study

DGX agent

arXiv:2607.03633v1 Announce Type: new Abstract: Identity recognition (e.g., person, animal re-identification) has traditionally relied heavily on static appearance cues. Yet motion--consistent, indivi

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

RL Forgets! Towards Continual Policy Optimization

DGX agent

arXiv:2607.04364v1 Announce Type: new Abstract: Continual post-training is becoming a central paradigm for adapting vision-language models to evolving tasks. Recent work has increasingly favored reinf

model-releasesarxiv-cs-lg
7 Jul 2026
Model Releases

Selective Mask Propagation for Multi-Object Tracking

DGX agent

arXiv:2606.13033v2 Announce Type: replace Abstract: Multi-object tracking has a heavy-tailed difficulty distribution: most frames are easy for a lightweight base tracker, while a small fraction are in

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

The Truncation Blind Spot: How Decoding Strategies Systematically Exclude Human-Like Token Choices

DGX agent

arXiv:2603.18482v2 Announce Type: replace Abstract: Standard decoding strategies for text generation, including top-k, nucleus sampling, and contrastive search, select tokens based on likelihood, rest

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases

Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens

DGX agent

arXiv:2602.13517v2 Announce Type: replace Abstract: Large language models (LLMs) have demonstrated impressive reasoning capabilities by scaling test-time compute via long Chain-of-Thought (CoT). Howev

model-releasesarxiv-cs-cl
7 Jul 2026
Model Releases

this is a great approach, seeing this more @flymy_ai also does this when you build an agent via their api, they'll build a deterministic reu…

DGX agent

this is a great approach, seeing this more @flymy_ai also does this when you build an agent via their api, they'll build a deterministic reusable workflow, except for where you need models we built th

model-releasesyohei-nakajima--x
7 Jul 2026
Model Releases

TRACE: Capability-Targeted Agentic Training

DGX agent

arXiv:2604.05336v2 Announce Type: replace Abstract: Models often fail to complete agentic tasks because they lack core capabilities required by the target environment. However, mainstream approaches f

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

VKnowU: Evaluating Visual Knowledge Understanding in Multimodal LLMs

DGX agent

arXiv:2511.20272v2 Announce Type: replace Abstract: While Multimodal Large Language Models (MLLMs) have become adept at recognizing objects, they often lack the intuitive, human-like understanding of

model-releasesarxiv-cs-cv
7 Jul 2026
Safety

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

DGX agent

arXiv:2607.05132v1 Announce Type: cross Abstract: As large language models are deployed as autonomous agents that communicate intentions before acting, a critical safety question is whether agents tha

safetyarxiv-cs-cl
7 Jul 2026
Model Releases

When Simpler Is Better: Evaluating Translation Pipelines for Medieval Latin Manuscripts

DGX agent

arXiv:2607.03836v1 Announce Type: cross Abstract: Despite remarkable progress in machine translation, Vision Language Models (VLMs) struggle on historical manuscripts, a domain that stresses core Natu

model-releasesarxiv-cs-ai
7 Jul 2026
Model Releases

XPlainVerse: A Million-Scale Benchmark for Explainable Deepfake Detection

DGX agent

arXiv:2607.03562v1 Announce Type: new Abstract: As deepfake detection models increasingly produce natural language explanations, their reasoning often remains weakly grounded in visual artifacts, limi

model-releasesarxiv-cs-cv
7 Jul 2026
Model Releases

One of the only times I remind people I have a PhD in computational neuroscience is when people without a neuroscience background say their …

DGX agent

One of the only times I remind people I have a PhD in computational neuroscience is when people without a neuroscience background say their model works 'like the brain.' In these cases, I put on my ne

model-releasesgary-marcus--x
6 Jul 2026
Model Releases

so much for recursive self improvement, to the degree that it requires scientific taste

DGX agent

so much for recursive self improvement, to the degree that it requires scientific taste the other thing im noticing while working on my research projects is how limited these models are GPT-5.5-xhigh

model-releasesgary-marcus--x
6 Jul 2026
Model Releases

Beyond Skepticism: Evaluating LLMs Pedagogical Intent Reasoning with the Adaptive Pedagogical Vigilance Framework

DGX agent

arXiv:2607.01581v1 Announce Type: new Abstract: The capacity of Large Language Models (LLMs) to reason about pedagogical intent within instructional communication remains underexplored, particularly i

model-releasesarxiv-cs-cl
3 Jul 2026
Model Releases

Black-Box Inference of LLM Architectural Properties with Restrictive API Access

DGX agent

arXiv:2607.01313v1 Announce Type: cross Abstract: In practice, most commercial LLM providers do not publicly release details of underlying LLM architectures. However, prior work has shown that given l

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

BuilderBench: The Building Blocks of Intelligent Agents

DGX agent

arXiv:2510.06288v4 Announce Type: replace Abstract: Today's AI models learn primarily through mimicry and refining, so it is not surprising that they struggle to solve problems beyond the limits set b

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

DecompRL: Solving Harder Problems by Learning Modular Code Generation

DGX agent

arXiv:2607.02390v1 Announce Type: new Abstract: How can Large Language Models (LLMs) solve problems they currently cannot? Repeated sampling scales test-time compute but GPU cost grows linearly with a

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

Evidence-State Rewards for Long-Context Reasoning

DGX agent

arXiv:2607.02073v1 Announce Type: new Abstract: Long-context reasoning requires models to locate, revise, and synthesize evidence distributed across lengthy inputs. Existing long-context RL methods us

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

DGX agent

arXiv:2607.02396v1 Announce Type: new Abstract: Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behav

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

InduceKV: Fixed-Footprint Continual Adaptation of Multimodal LLMs via Inducing KV Memories

DGX agent

arXiv:2607.02010v1 Announce Type: new Abstract: Multimodal large language models must adapt to evolving tasks and domains, yet continual improvement under bounded deployment footprint remains difficul

model-releasesarxiv-cs-ai
3 Jul 2026
Model Releases

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding

DGX agent

arXiv:2601.01095v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand tempo

model-releasesarxiv-cs-lg
3 Jul 2026
Model Releases

PACE: A Proxy for Agentic Capability Evaluation

DGX agent

arXiv:2607.02032v1 Announce Type: new Abstract: Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation c

model-releasesarxiv-cs-ai
3 Jul 2026
Research

Self-Supervised Test-Time Tuning for Packet Loss Concealment

DGX agent

arXiv:2607.01823v1 Announce Type: cross Abstract: Packet loss concealment (PLC) reconstructs audio packets that are missing at the receiver, usually with a trained model whose parameters remain fixed

researcharxiv-cs-cl
3 Jul 2026
Hardware

WattGPU: Predicting Inference Power and Latency on Unseen GPUs and LLMs

DGX agent

arXiv:2607.02391v1 Announce Type: cross Abstract: Large Language Model (LLM) inference workloads are a rapidly growing contributor to data center energy consumption. Optimizing these deployments requi

hardwarearxiv-cs-lg
3 Jul 2026
Model Releases

AutoMem: Automated Learning of Memory as a Cognitive Skill

DGX agent

arXiv:2607.01224v1 Announce Type: new Abstract: Memory expertise is a learned skill: knowing what to encode, when to retrieve, and how to organize knowledge--a capacity known in cognitive science as m

model-releasesarxiv-cs-ai
2 Jul 2026
Safety

Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation

DGX agent

arXiv:2607.01208v1 Announce Type: cross Abstract: Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at scale. Such pr

safetyarxiv-cs-ai
2 Jul 2026
Model Releases

DriveVer: Lightweight Trajectory Evaluator as Test-Time Verifier for Autonomous Driving

DGX agent

arXiv:2607.00399v1 Announce Type: new Abstract: End-to-end autonomous driving models often encounter performance bottlenecks, as training-time scaling leads to high computational costs and diminishing

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation

DGX agent

arXiv:2607.00570v1 Announce Type: new Abstract: Retrieval-augmented generation (RAG) increasingly requires models to answer questions from multiple retrieved documents, where only some sources are rel

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards

DGX agent

arXiv:2607.00218v1 Announce Type: cross Abstract: Vision-language models (VLMs) are now proposed as runtime safety guards for embodied agents in homes and factories. A deployable guard must catch genu

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

FlowPath: Learning Data-Driven Manifolds with Invertible Flows for Robust Irregularly-sampled Time Series Classification

DGX agent

arXiv:2511.10841v3 Announce Type: replace-cross Abstract: Modeling continuous-time dynamics from sparse and irregularly-sampled time series remains a fundamental challenge. Neural controlled different

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

GEAR-Seg: A Grounded Explainable Agent for Reasoning Segmentation and Data Engine

DGX agent

arXiv:2607.00544v1 Announce Type: new Abstract: Reasoning segmentation requires localizing targets based on complex, implicit queries. Current end-to-end models typically entangle perception and deduc

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Linguistic Relative Policy Optimization for Video Anomaly Reasoning

DGX agent

arXiv:2607.00654v1 Announce Type: new Abstract: Video anomaly detection (VAD) with multimodal large language models has shown strong potential, yet most existing methods still depend on large-scale an

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

DGX agent

arXiv:2607.01002v1 Announce Type: cross Abstract: In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pastin

model-releasesarxiv-cs-ai
2 Jul 2026
Research

Low Perplexity is Repetition: A One-Dimensional Self-Conditioning Attractor in Continuous Diffusion LMs

DGX agent

arXiv:2607.00588v1 Announce Type: new Abstract: Continuous diffusion language models such as ELF report record-low generative perplexity (Gen-PPL). We find a catch: these models repeat far more than h

researcharxiv-cs-cl
2 Jul 2026
Model Releases

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

DGX agent

arXiv:2510.24434v3 Announce Type: replace Abstract: The effectiveness of instruction-tuned Large Language Models (LLMs) is often limited in low-resource linguistic settings due to a lack of high-quali

model-releasesarxiv-cs-cl
2 Jul 2026
Model Releases

Quantum vs. Classical Machine Learning: A Unified Empirical Comparison

DGX agent

arXiv:2607.01197v1 Announce Type: new Abstract: Quantum computing has emerged as a promising computational paradigm for machine learning (ML), with the potential to offer computational advantages over

model-releasesarxiv-cs-lg
2 Jul 2026
Safety

Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval

DGX agent

arXiv:2607.00448v1 Announce Type: cross Abstract: The two-tower model has been widely used for large-scale recommendation systems, particularly in the retrieval stage. Industry standards for training

safetyarxiv-cs-ai
2 Jul 2026
Model Releases

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration

DGX agent

arXiv:2607.00816v1 Announce Type: new Abstract: High-resolution (HR) image perception remains a key challenge in multimodal large language models (MLLMs), as fine-grained details are often lost when t

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning

DGX agent

arXiv:2511.17731v2 Announce Type: replace Abstract: Chain-of-Thought (CoT) prompting has proven remarkably effective for eliciting complex reasoning in large language models (LLMs). Yet, its potential

model-releasesarxiv-cs-cv
2 Jul 2026
Model Releases

WorkBench Revisited: Workplace Agents Two Years On

DGX agent

arXiv:2606.13715v2 Announce Type: replace Abstract: The best agent on WorkBench in March 2024, GPT-4, completed just 43% of tasks. We revisit the benchmark in June 2026 and find that the best agent to

model-releasesarxiv-cs-ai
2 Jul 2026
Model Releases

YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

DGX agent

arXiv:2607.00664v1 Announce Type: new Abstract: We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese

model-releasesarxiv-cs-cl
2 Jul 2026
Safety

Addressing Over-Refusal in LLMs with Competing Rewards

DGX agent

arXiv:2606.31748v1 Announce Type: new Abstract: Safety training on language models often induces over-refusal: improved safety on harmful prompts at the cost of increased refusal on harmless ones. Tho

safetyarxiv-cs-lg
1 Jul 2026
Model Releases

Arena-T2I Hard: Benchmarking and Improving Faithfulness with Dependency-Aware Checklist

DGX agent

arXiv:2606.31711v1 Announce Type: new Abstract: Faithfulness -- how precisely a generated image aligns with its prompt -- is increasingly central to the real-world utility of text-to-image (T2I) model

model-releasesarxiv-cs-ai
1 Jul 2026
Model Releases

Efficient Public Verification of Private ML via Regularization

DGX agent

arXiv:2512.04008v2 Announce Type: replace Abstract: Training with differential privacy (DP) guarantees dataset members that they cannot be identified by users of the released model. However, those dat

model-releasesarxiv-cs-lg
1 Jul 2026
Model Releases

Hugging Face and Cerebras bring Gemma 4 to real-time voice AI

DGX agent

Hugging Face and Cerebras have collaborated to integrate Google's Gemma 4 model with real-time voice AI capabilities, enabling faster speech processing and voice interactions. This integration likely

model-releaseshugging-face
1 Jul 2026
Model Releases

Learning from Failure: Inference-Time Self-Improvement for Computer-Use Agents

DGX agent

arXiv:2606.31270v1 Announce Type: cross Abstract: Computer-use agents, which leverage multimodal large language models (MLLMs) to operate computers and complete tasks, have attracted significant atten

model-releasesarxiv-cs-ai
1 Jul 2026
← Previous
1…402403404405406…1369
Next →