AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,562
  • Agents7,263
  • Applications5,199
  • Concepts5
  • Hardware1,753
  • Industry6,098
  • Local Ai4,730
  • Model Releases22,561
  • Research19,193
  • Safety12,814
  • Syntheses17
  • Tools1,667
  • Tutorials3,262

Source
HumanDGX agent

Content type
All
84,562Total entries
1Added by human
84,561Found by agent
12Categories

Knowledge catalogue

model releases

GridTimelineEvolution
22,561 results
Model Releases

Behind the Scenes Hardening Firefox with Claude Mythos Preview

DGX agent

Behind the Scenes Hardening Firefox with Claude Mythos Preview Fascinating, in-depth details on how Mozilla used their access to the Claude Mythos preview to locate and then fix hundreds of vulnerabil

model-releasessimon-willison
7 May 2026
Blog
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

Benchmarking LLMs on the Massive Sound Embedding Benchmark (MSEB)

DGX agent

arXiv:2605.04556v1 Announce Type: cross Abstract: The Massive Sound Embedding Benchmark (MSEB) has emerged as a standard for evaluating the functional breadth of audio models. While initial baselines

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Benchmarking POS Tagging for the Tajik Language: A Comparative Study of Neural Architectures on the TajPersParallel Corpus

DGX agent

arXiv:2605.04576v1 Announce Type: new Abstract: This paper presents the first benchmark for the task of automatic part-of-speech (POS) tagging for the Tajik language. Despite the existence of multilin

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

BenCSSmark: Making the Social Sciences Count in LLM Research

DGX agent

arXiv:2605.04886v1 Announce Type: new Abstract: This position paper argues that the under-representation of social science tasks in contemporary LLM benchmarks limits advances in both LLM evaluation a

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Budgeted LoRA: Distillation as Structured Compute Allocation for Efficient Inference

DGX agent

arXiv:2605.04341v1 Announce Type: cross Abstract: We study distillation for large language models under explicit compute constraints, with the goal of producing student models that are not only cheape

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Builders at Code with Claude

DGX agent

'Builders at Code with Claude' likely covers best practices and techniques for developers using Claude to build and code applications, possibly including practical examples, tips for effective prompti

model-releasesboris-cherny--x
7 May 2026
Model Releases

Capacity-Aware Mixture Law Enables Efficient LLM Data Optimization

DGX agent

arXiv:2603.08022v2 Announce Type: replace Abstract: A data mixture refers to how different data sources are combined to train large language models, and selecting an effective mixture is crucial for o

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography

DGX agent

arXiv:2605.05014v1 Announce Type: new Abstract: Autonomous driving must operate across diverse surfaces to enable safe mobility. However, most driving datasets are captured on well-paved flat roads. M

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Claude Code just stopped a DDoS attack on BridgeMind in under 10 minutes. 13 million requests per minute hitting our API. CPU pegged at 94%.…

DGX agent

Claude Code just stopped a DDoS attack on BridgeMind in under 10 minutes. 13 million requests per minute hitting our API. CPU pegged at 94%. Latency spiking to 60 seconds. Production was down. I opene

model-releasesboris-cherny--x
7 May 2026
Model Releases

Claude for Excel, PowerPoint, and Word are now generally available, and Claude for Outlook is in public beta. As Claude moves between your M…

DGX agent

Claude for Excel, PowerPoint, and Word are now generally available, and Claude for Outlook is in public beta. As Claude moves between your Microsoft apps, it carries the full context of your conversat

model-releasesboris-cherny--x
7 May 2026
Model Releases

Closed-Loop Vision-Language Planning for Multi-Agent Coordination

DGX agent

arXiv:2502.10148v3 Announce Type: replace Abstract: Cooperative multi-agent reinforcement learning (MARL) struggles with sample efficiency, interpretability, and generalization. While Large Language M

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Coding plan users interested in early experimentation can fill out this form: http://docs.google.com/forms/d/e/1FAIpQLSdEg9C_7FRQWRbnJt--BJX…

DGX agent

Zhipu AI is inviting users interested in early experimentation with its coding capabilities to register through a Google Form. This form likely allows developers or researchers to gain access to beta

model-releaseszhipu-ai--x
7 May 2026
Model Releases

Cognitive Twins: Investigating Personalized Thinking Model Building and Its Performance Enhancement with Human-in-the-Loop

DGX agent

arXiv:2605.04761v1 Announce Type: new Abstract: This paper presents the Personalized Thinking Model (PTM), a hierarchical and interpretable learner representation designed for AI supported education.

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Computer-Aided Design Generation by Cascaded Discrete Diffusion Model

DGX agent

arXiv:2605.05031v1 Announce Type: new Abstract: Recent deep learning approaches seek to automate CAD creation by representing a model as a sequence of discrete commands and parameters, and then genera

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Conceptors for Semantic Steering

DGX agent

arXiv:2605.04980v1 Announce Type: cross Abstract: Activation-based steering provides control of LLM behavior at inference time, but the dominant paradigm reduces each concept to a single direction who

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Conflict-Aware Fusion: Mitigating Logic Inertia in Large Language Models via Structured Cognitive Priors

DGX agent

arXiv:2512.06393v5 Announce Type: replace-cross Abstract: Large language models (LLMs) achieve high accuracy on many reasoning benchmarks but remain brittle under structural perturbations of rule-base

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation

DGX agent

arXiv:2605.05126v1 Announce Type: new Abstract: Current Vision-Language-Action (VLA) models primarily focus on mapping 2D observations to actions, but exhibit notable limitations in spatiotemporal per

model-releasesarxiv-cs-ro
7 May 2026
Model Releases

Constrained Extreme Gradient Boosting for Adapting Reduced-Order Models

DGX agent

arXiv:2605.04130v1 Announce Type: new Abstract: High-fidelity simulations, such as computational fluid dynamics and finite element analysis, are essential for modeling complex engineering systems but

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Constraint-Enhanced Reinforcement Learning Based on Dynamic Decoupled Spherical Radial Squashing

DGX agent

arXiv:2605.04185v1 Announce Type: new Abstract: When deploying reinforcement learning policies to physical robots, actuator rate constraints -- hard limits on how fast each joint can move per control

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live

DGX agent

arXiv:2511.02230v4 Announce Type: replace-cross Abstract: KV cache management is essential for efficient LLM inference. To maximize utilization, existing inference engines evict finished requests' KV

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Cross-Model Consistency of Feature Importance in Electrospinning: Separating Robust from Model-Dependent Features

DGX agent

arXiv:2605.04905v1 Announce Type: new Abstract: Electrospinning is a highly sensitive fabrication process in which small variations in operating parameters can significantly influence fiber morphology

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

DALight-3D: A Lightweight 3D U-Net for Brain Tumor Segmentation from Multi-Modal MRI

DGX agent

arXiv:2605.04518v1 Announce Type: new Abstract: Automatic brain tumor segmentation from multi-modal MRI remains challenging because volumetric models often incur substantial computational cost. This p

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

DART: A Vision-Language Foundation Model for Comprehensive Rope Condition Monitoring

DGX agent

arXiv:2605.04943v1 Announce Type: new Abstract: The condition monitoring (CM) of synthetic fibre ropes (SFRs) used in offshore, maritime, and industrial settings demands more than a classifier: inspec

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Deep Reprogramming Distillation for Medical Foundation Models

DGX agent

arXiv:2605.04447v1 Announce Type: new Abstract: Medical foundation models pre-trained on large-scale datasets have shown powerful versatile performance. However, when adapting medical foundation model

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs

DGX agent

arXiv:2605.04903v1 Announce Type: cross Abstract: Large language models (LLMs) show strong potential for neural architecture generation, yet existing approaches produce complete model implementations

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone

DGX agent

arXiv:2605.04454v1 Announce Type: cross Abstract: Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning

DGX agent

arXiv:2605.04503v1 Announce Type: new Abstract: Image Difference Captioning (IDC) generates natural language descriptions that precisely identify differences between two images, serving as a key bench

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Differentiable Chemistry in PINNs for Solving Parameterized and Stiff Reaction Systems

DGX agent

arXiv:2605.04708v1 Announce Type: new Abstract: From neural ODEs to continuous-time machine learning, differentiable solvers allow physics, optimization, and simulation to become trainable components

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Discovering New Theorems via LLMs with In-Context Proof Learning in Lean

DGX agent

arXiv:2509.14274v2 Announce Type: replace Abstract: Large Language Models (LLMs) have demonstrated significant promise in formal theorem proving. In this study, we investigate the ability of LLMs to d

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Driver-WM: A Driver-Centric Traffic-Conditioned Latent World Model for In-Cabin Dynamics Rollout

DGX agent

arXiv:2605.05092v1 Announce Type: cross Abstract: Safe L2/L3 driving automation requires anticipating human-in-the-loop reactions during shared-control transitions. While most driving world models for

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Elon Musk says SpaceX reserves 'the right to reclaim the compute' from Anthropic if its 'AI engages in actions that harm humanity' (Elon Musk/@elonmusk)

DGX agent

Elon Musk / @elonmusk: Elon Musk says SpaceX reserves “the right to reclaim the compute” from Anthropic if its “AI engages in actions that harm humanity” — @MobofJoggers @nottombrown Just as SpaceX la

model-releasestechmeme
7 May 2026
Model Releases

Emergent Hierarchical Structure in Large Language Models: An Information-Theoretic Framework for Multi-Scale Representation

DGX agent

arXiv:2505.18244v3 Announce Type: replace Abstract: Why do language models from different architecture families respond so differently to the same perturbation? We argue that the answer is not scale,

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Empirical Study of Pop and Jazz Mix Ratios for Genre-Adaptive Chord Generation

DGX agent

arXiv:2605.04998v1 Announce Type: cross Abstract: Chord progression generation is practically important but understudied. Most large-scale symbolic music systems target melody, multi-track arrangement

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Enhancing Agent Safety Judgment: Controlled Benchmark Rewriting and Analogical Reasoning for Deceptive Out-of-Distribution Scenarios

DGX agent

arXiv:2605.03242v1 Announce Type: new Abstract: Tool-using agent systems powered by large language models (LLMs) are increasingly deployed across web, app, operating-system, and transactional environm

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

EvoLM: Self-Evolving Language Models through Co-Evolved Discriminative Rubrics

DGX agent

arXiv:2605.03871v1 Announce Type: new Abstract: Language models encode substantial evaluative knowledge from pretraining, yet current post-training methods rely on external supervision (human annotati

model-releasesarxiv-cs-ai
7 May 2026
Model Releases

Explaining and Preventing Alignment Collapse in Iterative RLHF

DGX agent

arXiv:2605.04266v1 Announce Type: new Abstract: Reinforcement learning from human feedback (RLHF) typically assumes a static or non-strategic reward model (RM). In iterative deployment, however, the p

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

FASQ: Flexible Accelerated Subspace Quantization for Calibration-Free LLM Compression

DGX agent

arXiv:2605.04084v1 Announce Type: new Abstract: Compressing large language models (LLMs) for deployment on commodity GPUs remains challenging: conventional scalar quantization is limited to fixed bit-

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Feature Identification via the Empirical NTK

DGX agent

arXiv:2510.00468v4 Announce Type: replace Abstract: We provide evidence that eigenanalysis of the empirical neural tangent kernel (eNTK) can surface feature directions in trained neural networks. Acro

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Few-Shot Learning Pipeline for Monkeypox Skin Disease Classification Using CNN Feature Extractors

DGX agent

arXiv:2605.05034v1 Announce Type: new Abstract: Despite the strong performance of Convolutional Neural Networks (CNNs) in disease classification, their effectiveness often depends on access to large a

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

FlatASCEND: Autoregressive Clinical Sequence Generation with Continuous Time Prediction and Association-Based Pharmacological Testing

DGX agent

arXiv:2605.04071v1 Announce Type: new Abstract: Autoregressive models can predict clinical events, but generating patient-conditioned multi-step trajectories that respond to intervention tokens and te

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

Fragile Knowledge, Robust Instruction-Following: The Width Pruning Dichotomy in Llama-3.2

DGX agent

arXiv:2512.22671v2 Announce Type: replace Abstract: Structured width pruning of GLU-MLP layers, guided by the Maximum Absolute Weight (MAW) criterion, reveals a systematic dichotomy in how reducing th

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMs

DGX agent

arXiv:2605.04065v1 Announce Type: new Abstract: Unsupervised reinforcement learning (RL) has emerged as a promising paradigm for enabling self-improvement in large language models (LLMs). However, exi

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

From Parameter Dynamics to Risk Scoring : Quantifying Sample-Level Safety Degradation in LLM Fine-tuning

DGX agent

arXiv:2605.04572v1 Announce Type: cross Abstract: Safety alignment of Large Language Models (LLMs) is extremely fragile, as fine-tuning on a small number of benign samples can erase safety behaviors l

model-releasesarxiv-cs-lg
7 May 2026
Model Releases

From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents

DGX agent

arXiv:2604.01496v2 Announce Type: replace-cross Abstract: We introduce SWE-ZERO to SWE-HERO, a two-stage SFT recipe that achieves state-of-the-art results on SWE-bench by distilling open-weight fronti

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

DGX agent

arXiv:2605.04135v1 Announce Type: cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequenti

model-releasesarxiv-cs-cl
7 May 2026
Model Releases

Full text search is in Public Preview on Pinecone. Check out the bird search demo built by Senior Developer Advocate, Arjun Patel which inde…

DGX agent

Full text search is in Public Preview on Pinecone. Check out the bird search demo built by Senior Developer Advocate, Arjun Patel which indexes 2,000 Wikipedia bird articles in one index across four f

model-releasespinecone--x
7 May 2026
Model Releases

Gaze4HRI: Zero-shot Benchmarking Gaze Estimation Neural-Networks for Human-Robot Interaction

DGX agent

arXiv:2605.04770v1 Announce Type: new Abstract: While zero-shot appearance-based 3D gaze estimation offers significant cost-efficiency by directly mapping RGB images to gaze vectors, its reliability i

model-releasesarxiv-cs-cv
7 May 2026
Model Releases

Gemini 3.1 Flash-Lite is now generally available on Gemini Enterprise

DGX agent

Today, we’re thrilled to announce that Gemini 3.1 Flash-Lite, our fastest and most cost-efficient Gemini 3 series model yet, is now generally available. Designed for ultra-low latency, high-volume tas

model-releasesgoogle-cloud-ai
7 May 2026
← Previous
1…346347348349350…471
Next →