AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries87,617
  • Agents7,497
  • Applications5,365
  • Concepts5
  • Hardware1,816
  • Industry6,151
  • Local Ai4,900
  • Model Releases23,593
  • Research19,967
  • Safety13,267
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries87,617
  • Agents7,497
  • Applications5,365
  • Concepts5
  • Hardware1,816
  • Industry6,151
  • Local Ai4,900
  • Model Releases23,593
  • Research19,967
  • Safety13,267
  • Syntheses17
  • Tools1,674
  • Tutorials3,365

Source
HumanDGX agent

87,617Total entries
1Added by human
87,616Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
62,987 results
10 Apr 2026

Possible memory leak in Ollama when using Claude Code?

Model ReleasesDGX agent

Users in the r/ollama community have reported a possible memory leak occurring in Ollama when it is used as a backend with Claude Code, with Ollama runner processes not always being properly termin...

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning

SafetyDGX agent

arXiv:2505.24499v2 Announce Type: replace Abstract: Generating high-quality Scalable Vector Graphics (SVGs) is challenging for Large Language Models (LLMs), as it requires advanced reasoning for struc

SAT: Balancing Reasoning Accuracy and Efficiency with Stepwise Adaptive Thinking

ResearchDGX agent

arXiv:2604.07922v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) have revolutionized complex problem-solving, yet they exhibit a pervasive 'overthinking', generating unnecessarily long

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SceneScribe-1M: A Large-Scale Video Dataset with Comprehensive Geometric and Semantic Annotations

Model ReleasesDGX agent

arXiv:2604.07990v1 Announce Type: new Abstract: The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both seman

Selective Neuron Amplification for Training-Free Task Enhancement

ResearchDGX agent

arXiv:2604.07098v1 Announce Type: new Abstract: Large language models often fail on tasks they seem to already understand. In our experiments, this appears to be less about missing knowledge and more

SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio

ApplicationsDGX agent

arXiv:2604.06389v1 Announce Type: new Abstract: Uncertainty estimation for reasoning language models remains difficult to deploy in practice: sampling-based methods are computationally expensive, whil

SentinelSphere: Integrating AI-Powered Real-Time Threat Detection with Cybersecurity Awareness Training

Model ReleasesDGX agent

arXiv:2604.06900v1 Announce Type: cross Abstract: The field of cybersecurity is confronted with two interrelated challenges: a worldwide deficit of qualified practitioners and ongoing human-factor wea

Should we be optimizing for limited compute instead of more parameters? Thoughts?

Local AiDGX agent

"The search did not return the specific Reddit thread. However, I can provide a summary based on what the topic is broadly about within the local-AI/Ollama community context:

TGIF! Here are some of our favorite updates from the past week: — Notebooks in @GeminiApp, an integration with @NotebookLM that enables you …

Model ReleasesDGX agent

TGIF! Here are some of our favorite updates from the past week: — Notebooks in @GeminiApp, an integration with @NotebookLM that enables you to retrieve context from your private notebooks or convert y

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training

SafetyDGX agent

arXiv:2604.07754v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) raises significant ethical and safety concerns. While LLM alignment techniques are adopted to improve m

The Geometry of Forgetting

Model ReleasesDGX agent

arXiv:2604.06222v1 Announce Type: cross Abstract: Why do we forget? Why do we remember things that never happened? The conventional answer points to biological hardware. We propose a different one: ge

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

Model ReleasesDGX agent

arXiv:2604.08401v1 Announce Type: cross Abstract: In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However

9 Apr 2026

Grok Law

Model ReleasesDGX agent

Grok Law Grok-4.20 just ranked #1 in Legal & Government on Chatbot Arena It’s officially outperforming Anthropic’s Opus 4.6 and Google’s Gemini 3.1 Pro Grok is actively helping people navigate real la

'Harnesses are intimately tied to memory, which means that by choosing an open harness you are choosing to own your memory, and not have it …

AgentsDGX agent

'Harnesses are intimately tied to memory, which means that by choosing an open harness you are choosing to own your memory, and not have it be locked into a proprietary harness or tied to a single mod

Introducing Gemma 4 31B from @GoogleDeepMind on Together AI. AI natives can now use Gemma 4 31B on Together and benefit from reliable infere…

Model ReleasesDGX agent

Introducing Gemma 4 31B from @GoogleDeepMind on Together AI. AI natives can now use Gemma 4 31B on Together and benefit from reliable inference for multimodal reasoning, tool use, and agentic workflow

Sparks unicorn https://x.com/emollick/status/2024756029121020236?s=20

Model ReleasesDGX agent

Sparks unicorn https://x.com/emollick/status/2024756029121020236?s=20 Here is the Gemini 3.1 'Sparks unicorn' (This is created using TikZ, which is a language built for scientific diagrams & very much

We should view the history of physics as a long-running program synthesis task. Kepler and Newton were searching the space of possible symbo…

ResearchDGX agent

We should view the history of physics as a long-running program synthesis task. Kepler and Newton were searching the space of possible symbolic models to find the simplest one that would best satisfy

8 Apr 2026

Again, if you care about computer security, read the red team report: https://red.anthropic.com/2026/mythos-preview/

SafetyDGX agent

Anthropic's Frontier Red Team report (April 2026) details the cybersecurity capabilities of Claude Mythos Preview, a general-purpose frontier model that performs strongly across the board but is s...

pay for opus to write slop code pay for mythos to fix slop code

IndustryDGX agent

Emad Mostaque (founder of Stability AI) posted a sardonic observation on X highlighting an ironic dynamic in the AI coding market: users pay for Claude Opus to generate low-quality 'slop' code, the...

this is interesting. 1. Did Anthropic forget to run a control? 2. Where does this leave us?

SafetyDGX agent

this is interesting. 1. Did Anthropic forget to run a control? 2. Where does this leave us? New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped anal

Today, we are launching our collaboration with @nomic_ai to make AI agents more effectively and efficiently understand complex PDF documents…

Local AiDGX agent

Today, we are launching our collaboration with @nomic_ai to make AI agents more effectively and efficiently understand complex PDF documents. Nomic's new nomic-layout-v1 model allows your AI agents to

Trying to DIY your own document parser by screenshotting into a frontier VLM (Opus, 5.4, Gemini) carries when you try to scale it up into pr…

Model ReleasesDGX agent

Trying to DIY your own document parser by screenshotting into a frontier VLM (Opus, 5.4, Gemini) carries when you try to scale it up into production workflows. Here are two edge cases we've observed:

Using a non-hermetic agent and having trouble with GPT 5.4 or your Codex sub? The rumors are true: it works a lot better with the Hermes spe…

AgentsDGX agent

Nous Research's **Hermes Agent** is an open-source agentic framework by NousResearch that delivers significantly improved reliability when using GPT-5.x or Codex (OpenAI Codex subscription) models ...

Writing fiction seems to be a genuine weak spot for LLMs that is not improving as rapidly as almost every other area. There may be a lot of …

Model ReleasesDGX agent

Writing fiction seems to be a genuine weak spot for LLMs that is not improving as rapidly as almost every other area. There may be a lot of reasons why this is happening. It would be a really interest

7 Apr 2026

Come check out GLM 5.1 in Code Arena for agentic web development tasks using tools. Don’t forget to vote, Code Arena scores are coming up ne…

AgentsDGX agent

GLM-5.1 is Zhipu AI's next-generation flagship model for agentic engineering, achieving state-of-the-art performance on SWE-Bench Pro and leading its predecessor GLM-5 by a wide margin on NL2Repo (...

“gpt2-large is too powerful to be publicly released” vibes

Model ReleasesDGX agent

Julien Chaumond (co-founder of Hugging Face) posted a tweet referencing the infamous 2019 OpenAI decision to initially withhold GPT-2-large from public release due to fears it was 'too dangerous,' ...

Thank you to @AnthropicAI for sending FFmpeg patches

Model ReleasesDGX agent

Thank you to @AnthropicAI for sending FFmpeg patches Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, C

20 Aug 2026

b10505

Model ReleasesDGX agent

server: add dedup-cache-models preset option (#27346) Website: https://llama.app macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED macOS Intel (x64) iOS XCF

ChildSafeAds Shared Task 2026: Commercial Content in Child-Facing YouTube Videos

Model ReleasesDGX agent

arXiv:2608.19165v1 Announce Type: new Abstract: ChildSafeAds is a shared task on commercial content in YouTube videos likely to reach children and teenagers. It contains 3,360 videos from 939 channels

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

Model ReleasesDGX agent

arXiv:2608.18734v1 Announce Type: new Abstract: 4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision e

Detecting Backdoors in Object Detection via Pre-NMS Prediction Distribution Shift

SafetyDGX agent

arXiv:2608.19088v1 Announce Type: cross Abstract: Object detection models deployed in safety-critical applications remain vulnerable to backdoor attacks that cause targeted misbehaviors when a hidden

Discretizing Continuous Time Series for Imputation with Masked Diffusion Training

ApplicationsDGX agent

arXiv:2608.19119v1 Announce Type: cross Abstract: Time series imputation is a crucial area for reliable time series analysis, yet it remains challenging due to the complex temporal dynamics and noise

FairGlucose: A CGM Fairness Benchmark Reveals Subgroup Disparities Hidden in Population-Level Validation

Model ReleasesDGX agent

arXiv:2608.18296v1 Announce Type: cross Abstract: As CGM-based AI tools approach clinical deployment, whether their accuracy is equitable across patient demographics remains insufficiently tested. To

From Storage to Access: Verifiable Activation of Parametric Knowledge in LLMs via Explicit Priming and Implicit Reasoning

SafetyDGX agent

arXiv:2608.18581v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) encode rich factual knowledge in their parameters, reliably recalling and verifying such knowledge remains a key

G9v3-39A5B on artificialanalysis looks good. Has anyone tested it?

Model ReleasesDGX agent

I see https://github.com/linuxid10t/llama.cpp/tree/feature/g9v3-support but yeah... Considering that some results place it above Qwen 3.6 27B (the top performer until a few days ago) and that it is an

Graphical Design of Interpretable Architectures

ResearchDGX agent

arXiv:2608.18936v1 Announce Type: cross Abstract: Designing, implementing, and comparing interpretable architectures requires a formal language to represent them. The most common representations fall

Mise-en-Scene: Implicit Layout Emergence in Diffusion Transformers for Human-AI Design Co-Creation

Model ReleasesDGX agent

arXiv:2608.19000v1 Announce Type: new Abstract: Automating graphic design synthesis from user-provided elements requires both a coherent overall composition and the exact preservation of each asset. E

MMD-Flagger: Leveraging Maximum Mean Discrepancy to Detect Hallucinations

Model ReleasesDGX agent

arXiv:2506.01367v4 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly integrated into agentic AI systems, yet their propensity to generate hallucinations remains a critical

Real-Time Control-Constrained DDP for Underactuated Balancing of Legged Robots

ResearchDGX agent

arXiv:2608.18552v1 Announce Type: new Abstract: This paper presents a real-time control-constrained Differential Dynamic Programming (DDP) framework for underactuated legged robots. To address the lim

Structure, Association, and Decision Value: Representation-Based Difficulty Estimation for Adaptive Inference in African-Language NLI

Model ReleasesDGX agent

arXiv:2608.19003v1 Announce Type: new Abstract: We ask whether internal representation statistics can provide useful example-level difficulty signals for adaptive inference in multilingual African NLP

Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering

Model ReleasesDGX agent

arXiv:2605.05962v2 Announce Type: replace Abstract: This paper addresses end-to-end geospatial question answering over multilingual toponymic data. We introduce a bilingual (Russian-Tatar) dataset of

19 Aug 2026

ArborMem: Navigating Interaction States with Memory Forests

Model ReleasesDGX agent

arXiv:2608.17534v1 Announce Type: new Abstract: Large language models increasingly serve as persistent conversational assistants, requiring memory that preserves relevant experience and maintains cont

Automated ACL Footprint Identification Using 3D Deep Learning

ResearchDGX agent

arXiv:2608.18012v1 Announce Type: new Abstract: One of the most common reasons for anterior cruciate ligament (ACL) reconstruction failure is femoral tunnel malpositioning (ACL footprint center and tu

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

SafetyDGX agent

arXiv:2503.20182v2 Announce Type: replace-cross Abstract: As large language models (LLMs) increasingly function as human-like assistants exhibiting human-like personality traits, understanding their b

ChannelFlow-Tools: A Configuration-Driven Pipeline for Generating Machine-Learning-Ready Datasets of 3D Obstructed Channel Flows

TutorialsDGX agent

arXiv:2509.15236v2 Announce Type: replace-cross Abstract: Data-driven surrogate models are increasingly used in computational fluid dynamics, and their reliability depends on the quality of the traini

Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL

AgentsDGX agent

arXiv:2608.17253v1 Announce Type: cross Abstract: Reinforcement learning (RL) has emerged as a powerful approach for improving reasoning in language and vision-language models, yet its strongest succe

ControlledShifts: Towards Standardizing Robustness Evaluation in Trajectory Prediction Under Distribution Shifts

Model ReleasesDGX agent

arXiv:2608.17882v1 Announce Type: new Abstract: Trajectory prediction is central to safety in autonomous driving, yet learning-based predictors tend to degrade sharply when encountering scenarios poor

DMT-Dens: Density-preserving manifold visualization for biological data

Model ReleasesDGX agent

arXiv:2608.17571v1 Announce Type: cross Abstract: Motivation: Low-dimensional embeddings are widely used to explore cell-state heterogeneity in single-cell and other high-dimensional biological data.

Efficient Resource Optimization for Split Federated Learning

ResearchDGX agent

arXiv:2608.17849v1 Announce Type: new Abstract: Split federated learning (SFL) has emerged as a powerful paradigm for model training at the edge. However, SFL inherently involves discrete decision var

Estimating Parameter Fields in Multi-Physics PDEs from Scarce Measurements

Model ReleasesDGX agent

arXiv:2509.00203v3 Announce Type: replace Abstract: Parameterized partial differential equations (PDEs) underpin the mathematical modeling of complex systems in diverse domains, including engineering,

EvoTS-Agent: A Self-Evolving LLM Agent for Financial Time Series Change Point Detection

Model ReleasesDGX agent

arXiv:2608.17933v1 Announce Type: new Abstract: Financial time series exhibit non-stationary and heterogeneous statistical properties, making change-point detection challenging because no single unsup

From Student Risk Prediction to SC2R: Semantics-Constrained Counterfactual Recourse for Educational Decision Support

ApplicationsDGX agent

arXiv:2608.17618v1 Announce Type: cross Abstract: Learning analytics models can identify students at risk of poor performance, but they do not directly indicate which interventions are feasible, actio

GLM-5.3 is out on AA, and I'm fed up with their Intelligence/cost plot

Model ReleasesDGX agent

I think AA's intelligence/cost plot is seriously misleading, so I decided to make my own. Their plot is in the second image. All points are at max thinking. All intelligence index scores are from AA.

Global Convergence of Gradient EM for Over-Parameterized Gaussian Mixtures

ResearchDGX agent

arXiv:2506.06584v2 Announce Type: replace Abstract: Learning Gaussian Mixture Models (GMMs) is a fundamental problem in statistics and machine learning, with the Expectation-Maximization (EM) algorith

Has anyone else noticed DeepSeek-V4-Flash:0731 using way more Ollama Cloud usage lately?

Model ReleasesDGX agent

I’ve been using deepseek-v4-flash:0731 on Ollama Cloud, and recently it feels like my usage is getting consumed much faster than before. My prompts and general workflow haven’t changed much, but the a

I benchmarked Qwen3.8 27B on browser tasks. It's on par with GPT 5.6 Luna (xhigh)

Model ReleasesDGX agent

The benchmark I used is BU bench v1, and the open source harness is Browser Agent. Qwen3.8 27B beat all other affordable or open models I tested. It's insanely good! submitted by /u/pierreb5 [link] [c

I might have found the perfect config parameters for qwen 3.8 27b

Model ReleasesDGX agent

Hello everyone, tried so hard to optimize my config and finally I simply get up to 70 t/s with q6 variant. And wanted to share with you guys so that other people with the same setup can enjoy. Please

Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries

Model ReleasesDGX agent

arXiv:2601.18899v3 Announce Type: replace-cross Abstract: Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking a f

Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation

Model ReleasesDGX agent

arXiv:2608.18041v1 Announce Type: new Abstract: Language has two parameters. Count how often words occur together and you estimate amplitude, the strength of association. Word embeddings and attention

Learning Where and What to Lift for Bi-planar X-ray-to-CT Reconstruction

Model ReleasesDGX agent

arXiv:2608.17255v1 Announce Type: cross Abstract: X-ray imaging can be approximately modeled as the projection of an underlying volumetric attenuation field, with each measurement recording the accumu

← Previous
1…372373374375376…1050
Next →