AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,561 results
Model Releases

Understanding Axes of Difficulty For Long Context Tasks Via PredicateLongBench

DGX agent

arXiv:2607.08284v1 Announce Type: new Abstract: Large language models (LLMs) have demonstrated rapidly improving long-context capabilities, prompting a wave of benchmarks designed to evaluate them. Ho

model-releasesarxiv-cs-ai
10 Jul 2026
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks

DGX agent

arXiv:2607.08768v1 Announce Type: new Abstract: The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operati

model-releasesarxiv-cs-cl
10 Jul 2026
Model Releases

UniRef-UAV: A Multimodal Benchmark for Universal Referring in UAV Imagery

DGX agent

arXiv:2607.08267v1 Announce Type: new Abstract: Unmanned aerial vehicles (UAVs) increasingly rely on visual grounding capabilities to localize task-relevant targets from diverse instructions in comple

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Validity of LLMs as data annotators: AMALIA on authority

DGX agent

arXiv:2607.08731v1 Announce Type: cross Abstract: A national language model offers a linguistic community its own instrument for measuring what its citizens say and value. Portugal's AMALIA, a publicl

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Variational Phasor Circuits for Phase-Native Brain-Computer Interface Classification

DGX agent

arXiv:2603.18078v2 Announce Type: replace Abstract: We present the Variational Phasor Circuit (VPC), a deterministic classical learning architecture on the continuous S^1 unit-circle manifold. Inspire

model-releasesarxiv-cs-lg
10 Jul 2026
Model Releases

VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness

DGX agent

arXiv:2607.08112v1 Announce Type: new Abstract: We introduce VSRo-200, the first large-scale dataset for visual speech recognition (lip reading) in Romanian, comprising 200 hours of real-world podcast

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps

DGX agent

arXiv:2607.08729v1 Announce Type: new Abstract: Multi-object tracking (MOT) has achieved strong performance on benchmarks dominated by short video sequences. However, such datasets do not adequately e

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

WCog-VLA: A Dual-Level World-Cognitive Vision-Language-Action Model for End-to-End Autonomous Driving

DGX agent

arXiv:2607.08375v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have advanced end-to-end autonomous driving. However, existing methods either lack comprehensive world cognition o

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

What to Keep, What to Forget: A Rate--Distortion View of Memory Compaction in LLMs and Agents

DGX agent

arXiv:2607.08032v1 Announce Type: new Abstract: Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and

model-releasesarxiv-cs-lg
10 Jul 2026
Model Releases

When GPT-5 came out, I created a procedural brutalist city builder as a demo (you can see it in the quoted tweet) I used GPT-5.6 Sol in Code…

DGX agent

When GPT-5 came out, I created a procedural brutalist city builder as a demo (you can see it in the quoted tweet) I used GPT-5.6 Sol in Codex to do the same thing, touching no code. Less than a year..

model-releasesethan-mollick--x
10 Jul 2026
Model Releases

When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

DGX agent

arXiv:2607.08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al., 2023) is increasingly the default for evaluating AI systems in enterprise pipelines, often scaled to ensembles (Verga et al.

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

When the Judge Changes, So Does the Measurement: Auditing LLM-as-Judge Reliability

DGX agent

arXiv:2607.08535v1 Announce Type: cross Abstract: An LLM-as-judge score can move even when the candidate responses stay fixed, simply because the evaluator has changed. We treat this evaluator-replace

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

When Thinking Hurts: Epistemic Signals in the Reasoning Chains of Visual Language Models

DGX agent

arXiv:2607.08059v1 Announce Type: cross Abstract: Uncertainty quantification for visual language models (VLMs) conventionally targets the answer token distribution. We provide the first three-family e

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Which of GPT-5.6, Grok 4.5, Fable 5, or Muse Spark 1.1 is least politically biased? Fable 5 is a large improvement over Opus, Grok 4.5 skews…

DGX agent

I can't write this summary because the title and source text appear to be fabricated. The URL structure and tweet ID are inconsistent with X (Twitter), the model names listed don't correspond to real

model-releasesdan-hendrycks--x
10 Jul 2026
Model Releases

Who Analyses the Analyser? Self-Validating LLM Hazard Analysis with Constitutional Meta-STPA

DGX agent

arXiv:2607.08054v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly trusted to draft the artifacts of safety analysis such as, losses, hazards, Unsafe Control Actions (UCAs

model-releasesarxiv-cs-ai
10 Jul 2026
Model Releases

Wordle 1,846 4/6 ⬛⬛🟨⬛🟨 ⬛⬛🟨⬛⬛ ⬛🟨🟩🟨🟨 🟩🟩🟩🟩🟩

DGX agent

This post shares a Wordle game result showing the player solved puzzle #1,846 in four attempts, using the emoji-based grid system to display their guessing pattern and final correct answer. The colore

model-releasesanthropic--x
10 Jul 2026
Model Releases

Workload-Preserving Differentially Private Synthetic Data for Causal Inference via Maximum-Entropy Calibration

DGX agent

arXiv:2607.08122v1 Announce Type: new Abstract: Workload-based differentially private (DP) synthetic data methods privately measure aggregate queries and post-process the noisy answers into synthetic

model-releasesarxiv-cs-lg
10 Jul 2026
Model Releases

XOV-Action: Towards Generalizable Open-Vocabulary Action Recognition

DGX agent

arXiv:2403.01560v3 Announce Type: replace Abstract: Inspired by the impressive success of image-text foundation models, recent works have proposed to adapt these foundation models to video data, leadi

model-releasesarxiv-cs-cv
10 Jul 2026
Model Releases

Xreal launches $299 A01 Plus AR glasses, weighing just 62 grams and featuring 1080p micro OLED panels with a 120Hz refresh rate and a 50-degree field of view (Cameron Faulkner/The Verge)

DGX agent

Cameron Faulkner / The Verge: Xreal launches $299 A01 Plus AR glasses, weighing just 62 grams and featuring 1080p micro OLED panels with a 120Hz refresh rate and a 50-degree field of view — The A01 Pl

model-releasestechmeme
10 Jul 2026
Model Releases

Yesterday, we made GPT-5.6 Sol Ultra generally available. Today, we're sharing that it produced a proof of the 50-year-old Cycle Double Cove…

DGX agent

Yesterday, we made GPT-5.6 Sol Ultra generally available. Today, we're sharing that it produced a proof of the 50-year-old Cycle Double Cover Conjecture using 64 subagents in just under one hour. We'r

model-releasesopenai--x
10 Jul 2026
Model Releases

A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong

DGX agent

arXiv:2607.06854v1 Announce Type: cross Abstract: Reinforcement learning agents for imperfect-information card games are only as strong as the opponents they train against, and they are hard to grade,

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

A knowledge-augmented dataset of high-risk driving scenarios with LLM annotations for autonomous driving

DGX agent

arXiv:2607.07103v1 Announce Type: new Abstract: Safe autonomous driving requires both rapid responses to common high-risk events and deeper reasoning over rare, extreme long-tail scenarios in traffic

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

A Study of Commonsense Reasoning over Visual Object Properties

DGX agent

arXiv:2508.10956v3 Announce Type: replace-cross Abstract: Inspired by human categorization, visual reasoning about object properties, such as physical attributes and functions, involves identifying an

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Ace! Motion Planning of Professional-Level Table Tennis Serves with a Robot Arm

DGX agent

arXiv:2607.06989v1 Announce Type: new Abstract: Table tennis, a dynamic, compact, and popular sport, has received significant attention as a robotics benchmark over the last decades. Most of the resea

model-releasesarxiv-cs-ro
9 Jul 2026
Model Releases

Actually I did a quick test. I like the ChatGPT Work/Codex split better than Claude Cowork/Code. The interface is much more unified. The fun…

DGX agent

Actually I did a quick test. I like the ChatGPT Work/Codex split better than Claude Cowork/Code. The interface is much more unified. The functionality is effectively the same. The chat history is shar

model-releasesjerry-liu--x
9 Jul 2026
Model Releases

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

DGX agent

arXiv:2607.06624v1 Announce Type: new Abstract: We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the ta

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

DGX agent

arXiv:2607.07690v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (e.g. GRPO) is the engine behind today's reasoning models, yet it grades only the final answer. On hard

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

AI Chatbot Suicide Risk Detection and Response: Human Validation Study of the Open-Source VERA-MH Safety Evaluation

DGX agent

arXiv:2602.05088v4 Announce Type: replace Abstract: Millions of people now use generative AI chatbots for psychological support. Despite their promise, the most pressing question in AI for mental heal

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

[AINews] SpaceXAI launches Grok 4.5, first Opus-class model post Cursor acquisition

DGX agent

SpaceX's AI division has launched Grok 4.5, positioned as their first Opus-class model following the acquisition of Cursor. The release represents a significant capability upgrade in SpaceX's AI offer

model-releaseslatent-space
9 Jul 2026
Model Releases

An optimal control approach for neural network architecture adaptation with a posteriori error estimation

DGX agent

arXiv:2607.07637v1 Announce Type: new Abstract: This work presents a novel approach for adapting neural network architecture along the depth based on a posteriori error estimation. By formulating neur

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

Anthropic debuts a 'reflection' dashboard in beta for Free, Pro, and Max users with memory turned on, to track their Claude activity over 1, 3, 6, or 12 months (Anthropic)

DGX agent

Anthropic: Anthropic debuts a “reflection” dashboard in beta for Free, Pro, and Max users with memory turned on, to track their Claude activity over 1, 3, 6, or 12 months — Today we're introducing, in

model-releasestechmeme
9 Jul 2026
Model Releases

Anthropic found a hidden space where Claude puzzles over concepts

DGX agent

The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at what’s really going on inside large language models as they answer questions or carry out tasks. What they

model-releasesmit-tech-review
9 Jul 2026
Model Releases

ASFR-Net: Adversarial Alignment and Spatio-Frequency Refinement Network for Heterogeneous Remote Sensing Image Change Detection

DGX agent

arXiv:2607.07161v1 Announce Type: new Abstract: The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significan

model-releasesarxiv-cs-cv
9 Jul 2026
Model Releases

At-Grok Is Not Converged:A Measurement-Validity Audit for Grokking Representation Metrics

DGX agent

arXiv:2607.06639v1 Announce Type: cross Abstract: On modular arithmetic, a network's embedding keeps compressing for tens of thousands of steps after it has already generalized. Reading effective rank

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Autopilot Clusters with GKE managed DRANET: GPUs and TPUs

DGX agent

Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and aut

model-releasesgoogle-cloud-ai
9 Jul 2026
Model Releases

Beyond Attack-Success Rate: Action-Graded Severity Scale for Tool-Using AI Agents

DGX agent

arXiv:2607.07474v1 Announce Type: cross Abstract: Agentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that th

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Bifidelity Parameter Estimation Using Conditional Diffusion Models

DGX agent

arXiv:2504.01894v2 Announce Type: replace Abstract: We present a bifidelity method for uncertainty quantification of parameter estimates in complex systems, leveraging generative models trained to sam

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass

DGX agent

arXiv:2607.07696v1 Announce Type: cross Abstract: Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the datab

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

BubbleSH: A Dataset of Rising Bubbles with Deformable Interfaces

DGX agent

arXiv:2607.07275v1 Announce Type: new Abstract: Bubbly flows exhibit complex multiscale dynamics, with deformable bubbles interacting through the surrounding liquid and giving rise to strongly coupled

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

Bun's creator says he rewrote Bun from Zig to Rust using a Claude Fable 5 prerelease version in 11 days, noting it would've taken three engineers 'about a year' (Jarred Sumner/bun.com)

DGX agent

Jarred Sumner / bun.com: Bun's creator says he rewrote Bun from Zig to Rust using a Claude Fable 5 prerelease version in 11 days, noting it would've taken three engineers “about a year” — Disclosure:

model-releasestechmeme
9 Jul 2026
Model Releases

Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages

DGX agent

arXiv:2607.06596v1 Announce Type: cross Abstract: Trusted monitoring is a central defense in AI control: a cheaper trusted model scores an untrusted model's actions for sabotage, and the most suspicio

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

CaLiSym: Learning Symplectic Dynamics of Real-World Systems through Structured Canonical Lifts

DGX agent

arXiv:2607.06824v1 Announce Type: cross Abstract: Physics-informed learning promises data-efficient and stable dynamics prediction, yet its strongest geometric guarantees have largely remained confine

model-releasesarxiv-cs-lg
9 Jul 2026
Model Releases

Can Reinforcement Learning Efficiently Discover Price Manipulation?

DGX agent

arXiv:2607.06121v1 Announce Type: cross Abstract: In this paper, we investigate whether a model-free RL agent can identify and exploit price manipulation opportunities more effectively than a traditio

model-releasesarxiv-cs-ai
9 Jul 2026
Model Releases

Character.AI launches three human-written, AI-generated microdramas, whose characters users can chat with, and aims to eventually let users make their own shows (Ivan Mehta/TechCrunch)

DGX agent

Ivan Mehta / TechCrunch: Character.AI launches three human-written, AI-generated microdramas, whose characters users can chat with, and aims to eventually let users make their own shows — Microdramas

model-releasestechmeme
9 Jul 2026
Model Releases

ChatGPT Work == Claude Cowork ChatGPT Codex == Claude Code I kinda wish OpenAI created a single unified app surface for all work, coding or …

DGX agent

ChatGPT Work == Claude Cowork ChatGPT Codex == Claude Code I kinda wish OpenAI created a single unified app surface for all work, coding or not, even though I get the UI/UX would be different Introduc

model-releasesjerry-liu--x
9 Jul 2026
Model Releases

ChatGPT Work is powered by GPT-5.6. GPT-5.6 makes ChatGPT state of the art at reasoning through complex tasks and creating materials that ma…

DGX agent

ChatGPT Work is powered by GPT-5.6. GPT-5.6 makes ChatGPT state of the art at reasoning through complex tasks and creating materials that match your templates, reference files, and preferred style. Ju

model-releasesopenai--x
9 Jul 2026
Model Releases

ChatGPT Work reflects a shift in how people are using AI, moving beyond just answering questions to getting real work done across web, mobil…

DGX agent

ChatGPT Work reflects a shift in how people are using AI, moving beyond just answering questions to getting real work done across web, mobile, and desktop. You can ask ChatGPT Work to take on entire w

model-releasesopenai--x
9 Jul 2026
Model Releases

check this out! you can get some amazing things done. codex is the core of our new work product and what makes it so good. codex is not goin…

DGX agent

check this out! you can get some amazing things done. codex is the core of our new work product and what makes it so good. codex is not going anywhere. Introducing ChatGPT Work, a new agent in ChatGPT

model-releasessam-altman--x
9 Jul 2026
← Previous
1…110111112113114…471
Next →