AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries90,914
  • Agents7,747
  • Applications5,533
  • Concepts5
  • Hardware1,920
  • Industry6,196
  • Local Ai5,094
  • Model Releases24,727
  • Research20,781
  • Safety13,739
  • Syntheses17
  • Tools1,678
  • Tutorials3,477

Source
HumanDGX agent

Content type
90,914Total entries
1Added by human
90,913Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,675 results
Model Releases

any reasonably fast public benchmarks I should run quants of deepseek flash 0731 on?

DGX agent

I have various quants of this model and am curious how they perform. can anyone recommend which benchmark would be a good test case for quantization effects? Maybe that can be completed with about 1 m

model-releasesr-localllama
8 Aug 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Anyone else amped up over Qwen 3.8?

DGX agent

I’ve been using 3.6 27B Q4, and that quant is fast on an M5. The code has been average, but consistently “good enough.” And, after a year, I can see home LLMs being served at home much like streaming

model-releasesr-localllama
8 Aug 2026
Model Releases

Auto mode is now the default in Claude Code for Pro, Max, and Team plans

DGX agent

Auto mode is now the default in Claude Code for Pro, Max, and Team plans Anthropic are really confident in Claude Code's auto mode, to the point that they are making it the default setting for new ses

model-releasessimon-willison
8 Aug 2026
Model Releases

Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected

DGX agent

I compared Qwen 35B-A3B MoE against Qwen 27B dense on a series of local coding-maintenance tasks. On my R9700/llama.cpp setup, the MoE model generated about 3.9× faster (~116 vs ~30 tok/s), but the co

model-releasesr-localllama
8 Aug 2026
Model Releases

Tesla V100 Qwen3.6 27B Performance

DGX agent

Looking for V100 users to share your config and it's performance. GPU: Tesla V100 PCIE 32Gb Qwen3.6 27B Q4_K_M + Q8_0 MTP 128K context length Pi coding agent llama.cpp model preset: [*] spec-default =

model-releasesr-localllama
8 Aug 2026
Model Releases

Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation

DGX agent

arXiv:2503.03556v3 Announce Type: replace Abstract: Object affordance reasoning, the ability to infer object functionalities based on physical properties, is fundamental for task-oriented planning and

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasks

DGX agent

arXiv:2608.05266v1 Announce Type: new Abstract: Large language model agents are increasingly being developed to control a wide range of scientific characterization tools including microscopes and sync

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Anyone running DeepSeek-V4-Flash-0731 on MI325X with vLLM? Mine is behaving completely broken

DGX agent

Is anyone here successfully running DeepSeek-V4-Flash-0731 locally with vLLM, especially on AMD MI325X? My setup: GPU: 1x AMD Instinct MI325X Model: deepseek-ai/DeepSeek-V4-Flash-0731 vLLM: 0.26.0 ROC

model-releasesr-localllama
7 Aug 2026
Model Releases

Beyond Sequence Order: Syntax-Informed Positional Embeddings for Transformers

DGX agent

arXiv:2608.06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to extit{syntactic structure}. We introduce extbf{

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Continual Learning in Transition

DGX agent

arXiv:2608.06216v1 Announce Type: cross Abstract: Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g.,

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

GROM: Gradient-Free Rapid One-Shot Machine Unlearning

DGX agent

arXiv:2608.05783v1 Announce Type: cross Abstract: Machine unlearning has become a critical capability for safely removing specific, sensitive knowledge from large language models (LLMs). Current state

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents

DGX agent

arXiv:2606.21140v2 Announce Type: replace-cross Abstract: Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs d

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

DGX agent

arXiv:2608.05539v1 Announce Type: new Abstract: Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D obj

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

DGX agent

arXiv:2604.04444v2 Announce Type: replace Abstract: Open-vocabulary object detection (OVOD) enables models to detect any object category, including unseen ones. Benefiting from large-scale pre-trainin

model-releasesarxiv-cs-cv
7 Aug 2026
Safety

Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training

DGX agent

arXiv:2608.06125v1 Announce Type: new Abstract: Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human prefer

safetyarxiv-cs-cv
7 Aug 2026
Model Releases

SEAM: Global consistency beyond local accuracy in scientific machine learning

DGX agent

arXiv:2608.05702v1 Announce Type: new Abstract: Scientific machine learning commonly validates models at the level of a subdomain, a benchmark split, or an explanation for one prediction. Yet such loc

model-releasesarxiv-cs-lg
7 Aug 2026
Model Releases

SemiAdapt-Instruct: Extensible Instruction Tuning via Latent Domain-Specialised Adapters

DGX agent

arXiv:2608.05161v1 Announce Type: new Abstract: Instruction-tuned LLMs are deployed into environments where domains evolve, yet extending a fine-tuned model's capabilities without full retraining rema

model-releasesarxiv-cs-cl
7 Aug 2026
Research

Timestep-Conditioned Transformers for Global Weather Forecasting

DGX agent

arXiv:2608.06241v1 Announce Type: new Abstract: Existing machine-learning weather forecasting models rely on predetermined and fixed autoregressive timesteps. The choice of model timestep involves a f

researcharxiv-cs-lg
7 Aug 2026
Model Releases

Tree-NET: Enhancing 2D Medical Image Segmentation Through Efficient Low-Level Feature Training

DGX agent

arXiv:2501.02140v2 Announce Type: replace-cross Abstract: This paper introduces Tree-NET, a novel framework for medical image segmentation that leverages bottleneck supervision to enhance both segment

model-releasesarxiv-cs-cv
7 Aug 2026
Model Releases

Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders

DGX agent

arXiv:2608.04586v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved significant success in speech-to-text translation (S2TT). However, when processing multilingual

model-releasesarxiv-cs-ai
6 Aug 2026
Safety

critical history and context that a lot of people have conveniently forgotten

DGX agent

critical history and context that a lot of people have conveniently forgotten For a very long time most high-performing AI models were end-to-end neural models; vector input -> vector output, with onl

safetygary-marcus--x
6 Aug 2026
Model Releases

Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs

DGX agent

arXiv:2608.04488v1 Announce Type: new Abstract: Despite rapid advances in large language models (LLMs), deploying and personalizing them on resource-constrained devices remains impractical due to high

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

Foreseeing the Invisible: Amodal Reconstruction of Leaf Fossil Images

DGX agent

arXiv:2608.04423v1 Announce Type: new Abstract: Fossil leaves are rarely preserved whole -- sedimentary rock hides, breaks, and erodes the lamina, yet paleobotany depends on the complete shape and out

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Gradient Immunity: Null-Space Resistance to Malicious Fine-Tuning

DGX agent

arXiv:2608.05045v1 Announce Type: cross Abstract: Released aligned large language models remain vulnerable to malicious downstream finetuning. Existing defenses are largely designed for the fine-tunin

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

How come artificialanalysis.ai ranks Gemma4 above Qwen3.6 27b in SciCode

DGX agent

Just came across this coding benchmark: SciCode Artificialanalysis.ai reports a ranking which contradicts the feeling we've towards those models in real life coding. Is Gemma 4 really that good, or a

model-releasesr-localllama
6 Aug 2026
Model Releases

Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness

DGX agent

arXiv:2608.04519v1 Announce Type: new Abstract: Benchmarking machine unlearning methods is critical to understand whether sensitive knowledge is removed from large language models (LLMs) or not. Curre

model-releasesarxiv-cs-ai
6 Aug 2026
Safety

Looking in the Mirror: Introspecting Side-Effect Misalignments Induced by Fine-Tuning

DGX agent

arXiv:2608.04347v1 Announce Type: new Abstract: Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competen

safetyarxiv-cs-lg
6 Aug 2026
Model Releases

MOON3.0: Reasoning-aware Multimodal Representation Learning for E-commerce Product Understanding

DGX agent

arXiv:2604.00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention. Alth

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

On the Effectiveness of Adaptation Strategies for VLM-Based Federated Learning in Remote Sensing

DGX agent

arXiv:2608.04791v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative training of deep learning models across decentralized image archives without requiring data centralization

model-releasesarxiv-cs-cv
6 Aug 2026
Agents

Outlook where frontier AI is headed next 18 months: The AI reasoning training + harness loop works if you can produce enough data and reason…

DGX agent

Outlook where frontier AI is headed next 18 months: The AI reasoning training + harness loop works if you can produce enough data and reasoning traces (via verifiers). Proven with code and math result

agentsfrancois-chollet--x
6 Aug 2026
Model Releases

Protoreasoning in Tiny Transformers

DGX agent

arXiv:2608.04980v1 Announce Type: cross Abstract: We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-ste

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

DGX agent

arXiv:2608.04385v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely incre

model-releasesarxiv-cs-cv
6 Aug 2026
Research

Right Reset: Chunking by Prefix Removal

DGX agent

arXiv:2608.04330v1 Announce Type: new Abstract: Removing the left context from a causal language model reveals a useful kind of boundary: an edge where the model processes the same right-hand tokens w

researcharxiv-cs-cl
6 Aug 2026
Model Releases

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?

DGX agent

arXiv:2608.04828v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

The Order Is the Guarantee: Verifier-Budgeted Code Deletion with Static-First Learned Proposals

DGX agent

arXiv:2608.04611v1 Announce Type: cross Abstract: Frontier coding models now match or exceed strong human reference points on programming benchmarks, yet benchmark success does not imply maintainable

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Thinking with Anchors: Grounded and Efficient Document Reasoning

DGX agent

arXiv:2608.04424v1 Announce Type: new Abstract: Existing document understanding benchmarks have largely focused on locating page elements, yet real-world document intelligence requires models to reaso

model-releasesarxiv-cs-cv
6 Aug 2026
Model Releases

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

DGX agent

arXiv:2608.03569v1 Announce Type: new Abstract: Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in eval

model-releasesarxiv-cs-ai
5 Aug 2026
Agents

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?

DGX agent

arXiv:2608.00155v1 Announce Type: cross Abstract: Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominan

agentsarxiv-cs-lg
5 Aug 2026
Model Releases

Approximate Speculative Decoding

DGX agent

arXiv:2608.03447v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel. Under standard greedy verificat

model-releasesarxiv-cs-ai
5 Aug 2026
Research

Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling

DGX agent

arXiv:2608.02618v1 Announce Type: new Abstract: Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized con

researcharxiv-cs-ai
5 Aug 2026
Safety

CT-HEG: A Bidirectional, Timestamp-Attributed Event Graph for ICU In-Hospital Mortality Prediction - An Architectural Ablation Study

DGX agent

arXiv:2608.02663v1 Announce Type: cross Abstract: Accurate ICU mortality prediction requires modeling irregular clinical observations across heterogeneous entity types. Existing sequence models handle

safetyarxiv-cs-ai
5 Aug 2026
Model Releases

DiagChain: A Diagnostic Benchmark for Evaluating LLM Agents on Evidence-Grounded Attack Chain Reconstruction

DGX agent

arXiv:2608.03591v1 Announce Type: cross Abstract: Large Language Model (LLM) agents offer a promising approach to attack chain reconstruction by retrieving and interpreting heterogeneous telemetry to

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Disentangling MLP Neuron Weights in Vocabulary Space

DGX agent

arXiv:2604.06005v2 Announce Type: replace Abstract: Interpreting the information encoded in language model weights remains a fundamental challenge in mechanistic interpretability. In this work, we int

model-releasesarxiv-cs-cl
5 Aug 2026
Model Releases

Distractor-Aware Truncation: Disentangling Context-Length Effects from Signal Loss in Long-Context LLM Benchmarks

DGX agent

arXiv:2608.03297v1 Announce Type: new Abstract: A standard claim in the literature on retrieval-augmented and memory-augmented language models is that shorter context is better when the relevant infor

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Don't Walk the Line: Boundary Guidance for Filtered Generation

DGX agent

arXiv:2510.11834v3 Announce Type: replace-cross Abstract: Generative models are increasingly paired with safety classifiers that filter harmful or undesirable outputs. A common strategy is to fine-tun

model-releasesarxiv-cs-cl
5 Aug 2026
Applications

Enactive Artificial Intelligence: A Decision-Centric Architecture for Complex Systems

DGX agent

arXiv:2608.03413v1 Announce Type: new Abstract: As artificial intelligence (AI) continues to evolve and mature, recent AI practices have moved beyond large language models (LLMs) and text or image gen

applicationsarxiv-cs-ai
5 Aug 2026
Model Releases

Evaluating LLMs in Database Scenarios: A Lifecycle Benchmark for Assessing Their Potential in Core Database Tasks

DGX agent

arXiv:2608.03794v1 Announce Type: cross Abstract: Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administra

model-releasesarxiv-cs-ai
5 Aug 2026
Model Releases

Externally Validated Breast Ultrasound Segmentation via Multi-task Learning with BI-RADS-Consistent Morphological Priors

DGX agent

arXiv:2511.15968v2 Announce Type: replace-cross Abstract: External validation of breast ultrasound segmentation models remains limited because internal train--test splits do not capture domain shifts

model-releasesarxiv-cs-ai
5 Aug 2026
← Previous
1…395396397398399…1369
Next →