AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries91,060
  • Agents7,763
  • Applications5,542
  • Concepts5
  • Hardware1,932
  • Industry6,210
  • Local Ai5,103
  • Model Releases24,798
  • Research20,784
  • Safety13,745
  • Syntheses17
  • Tools1,680
  • Tutorials3,481

Source
HumanDGX agent

Content type
91,060Total entries
1Added by human
91,059Found by agent
12Categories

Knowledge catalogue

Search: “models”

GridTimelineEvolution
65,793 results
Research

Selective Neuron Amplification for Training-Free Task Enhancement

DGX agent

arXiv:2604.07098v1 Announce Type: new Abstract: Large language models often fail on tasks they seem to already understand. In our experiments, this appears to be less about missing knowledge and more

researcharxiv-cs-lg
10 Apr 2026
Applications
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

SELFDOUBT: Uncertainty Quantification for Reasoning LLMs via the Hedge-to-Verify Ratio

DGX agent

arXiv:2604.06389v1 Announce Type: new Abstract: Uncertainty estimation for reasoning language models remains difficult to deploy in practice: sampling-based methods are computationally expensive, whil

applicationsarxiv-cs-ai
10 Apr 2026
Model Releases

SentinelSphere: Integrating AI-Powered Real-Time Threat Detection with Cybersecurity Awareness Training

DGX agent

arXiv:2604.06900v1 Announce Type: cross Abstract: The field of cybersecurity is confronted with two interrelated challenges: a worldwide deficit of qualified practitioners and ongoing human-factor wea

model-releasesarxiv-cs-ai
10 Apr 2026
Local Ai

Should we be optimizing for limited compute instead of more parameters? Thoughts?

DGX agent

"The search did not return the specific Reddit thread. However, I can provide a summary based on what the topic is broadly about within the local-AI/Ollama community context:

local-air-ollama
10 Apr 2026
Model Releases

TGIF! Here are some of our favorite updates from the past week: — Notebooks in @GeminiApp, an integration with @NotebookLM that enables you …

DGX agent

TGIF! Here are some of our favorite updates from the past week: — Notebooks in @GeminiApp, an integration with @NotebookLM that enables you to retrieve context from your private notebooks or convert y

model-releasesgoogle-ai--x
10 Apr 2026
Safety

The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training

DGX agent

arXiv:2604.07754v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) raises significant ethical and safety concerns. While LLM alignment techniques are adopted to improve m

safetyarxiv-cs-cl
10 Apr 2026
Model Releases

The Geometry of Forgetting

DGX agent

arXiv:2604.06222v1 Announce Type: cross Abstract: Why do we forget? Why do we remember things that never happened? The conventional answer points to biological hardware. We propose a different one: ge

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Verify Before You Commit: Towards Faithful Reasoning in LLM Agents via Self-Auditing

DGX agent

arXiv:2604.08401v1 Announce Type: cross Abstract: In large language model (LLM) agents, reasoning trajectories are treated as reliable internal beliefs for guiding actions and updating memory. However

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Grok Law

DGX agent

Grok Law Grok-4.20 just ranked #1 in Legal & Government on Chatbot Arena It’s officially outperforming Anthropic’s Opus 4.6 and Google’s Gemini 3.1 Pro Grok is actively helping people navigate real la

model-releaseselon-musk--x
9 Apr 2026
Agents

'Harnesses are intimately tied to memory, which means that by choosing an open harness you are choosing to own your memory, and not have it …

DGX agent

'Harnesses are intimately tied to memory, which means that by choosing an open harness you are choosing to own your memory, and not have it be locked into a proprietary harness or tied to a single mod

agentsharrison-chase--x
9 Apr 2026
Model Releases

Introducing Gemma 4 31B from @GoogleDeepMind on Together AI. AI natives can now use Gemma 4 31B on Together and benefit from reliable infere…

DGX agent

Introducing Gemma 4 31B from @GoogleDeepMind on Together AI. AI natives can now use Gemma 4 31B on Together and benefit from reliable inference for multimodal reasoning, tool use, and agentic workflow

model-releasestogether-ai--x
9 Apr 2026
Model Releases

Sparks unicorn https://x.com/emollick/status/2024756029121020236?s=20

DGX agent

Sparks unicorn https://x.com/emollick/status/2024756029121020236?s=20 Here is the Gemini 3.1 'Sparks unicorn' (This is created using TikZ, which is a language built for scientific diagrams & very much

model-releasesethan-mollick--x
9 Apr 2026
Research

We should view the history of physics as a long-running program synthesis task. Kepler and Newton were searching the space of possible symbo…

DGX agent

We should view the history of physics as a long-running program synthesis task. Kepler and Newton were searching the space of possible symbolic models to find the simplest one that would best satisfy

researchfrancois-chollet--x
9 Apr 2026
Safety

Again, if you care about computer security, read the red team report: https://red.anthropic.com/2026/mythos-preview/

DGX agent

Anthropic's Frontier Red Team report (April 2026) details the cybersecurity capabilities of Claude Mythos Preview, a general-purpose frontier model that performs strongly across the board but is s...

safetyethan-mollick--x
8 Apr 2026
Industry

pay for opus to write slop code pay for mythos to fix slop code

DGX agent

Emad Mostaque (founder of Stability AI) posted a sardonic observation on X highlighting an ironic dynamic in the AI coding market: users pay for Claude Opus to generate low-quality 'slop' code, the...

industryemad-mostaque--x
8 Apr 2026
Safety

this is interesting. 1. Did Anthropic forget to run a control? 2. Where does this leave us?

DGX agent

this is interesting. 1. Did Anthropic forget to run a control? 2. Where does this leave us? New post: We tested the Mythos showcase vulnerabilities with open models. They recovered similar scoped anal

safetygary-marcus--x
8 Apr 2026
Local Ai

Today, we are launching our collaboration with @nomic_ai to make AI agents more effectively and efficiently understand complex PDF documents…

DGX agent

Today, we are launching our collaboration with @nomic_ai to make AI agents more effectively and efficiently understand complex PDF documents. Nomic's new nomic-layout-v1 model allows your AI agents to

local-ainomic-ai--x
8 Apr 2026
Model Releases

Trying to DIY your own document parser by screenshotting into a frontier VLM (Opus, 5.4, Gemini) carries when you try to scale it up into pr…

DGX agent

Trying to DIY your own document parser by screenshotting into a frontier VLM (Opus, 5.4, Gemini) carries when you try to scale it up into production workflows. Here are two edge cases we've observed:

model-releasesjerry-liu--x
8 Apr 2026
Agents

Using a non-hermetic agent and having trouble with GPT 5.4 or your Codex sub? The rumors are true: it works a lot better with the Hermes spe…

DGX agent

Nous Research's **Hermes Agent** is an open-source agentic framework by NousResearch that delivers significantly improved reliability when using GPT-5.x or Codex (OpenAI Codex subscription) models ...

agentsnous-research--x
8 Apr 2026
Model Releases

Writing fiction seems to be a genuine weak spot for LLMs that is not improving as rapidly as almost every other area. There may be a lot of …

DGX agent

Writing fiction seems to be a genuine weak spot for LLMs that is not improving as rapidly as almost every other area. There may be a lot of reasons why this is happening. It would be a really interest

model-releasesethan-mollick--x
8 Apr 2026
Agents

Come check out GLM 5.1 in Code Arena for agentic web development tasks using tools. Don’t forget to vote, Code Arena scores are coming up ne…

DGX agent

GLM-5.1 is Zhipu AI's next-generation flagship model for agentic engineering, achieving state-of-the-art performance on SWE-Bench Pro and leading its predecessor GLM-5 by a wide margin on NL2Repo (...

agentszhipu-ai--x
7 Apr 2026
Model Releases

“gpt2-large is too powerful to be publicly released” vibes

DGX agent

Julien Chaumond (co-founder of Hugging Face) posted a tweet referencing the infamous 2019 OpenAI decision to initially withhold GPT-2-large from public release due to fears it was 'too dangerous,' ...

model-releasesyann-lecun--x
7 Apr 2026
Model Releases

Thank you to @AnthropicAI for sending FFmpeg patches

DGX agent

Thank you to @AnthropicAI for sending FFmpeg patches Introducing Project Glasswing: an urgent initiative to help secure the world’s most critical software. It’s powered by our newest frontier model, C

model-releasesboris-cherny--x
7 Apr 2026
Model Releases

A Hybrid Two-Stage Machine Learning Pipeline for Fault Detection and Classification in Power Transmission Systems

DGX agent

arXiv:2608.23726v1 Announce Type: cross Abstract: Rapid and accurate fault detection in high-voltage transmission networks is essential for grid reliability and equipment protection. Transmission faul

model-releasesarxiv-cs-lg
26 Aug 2026
Model Releases

Big shoutout to @perplexity_ai for properly benchmarking document understanding in their Portable Computer release 🔥 They took a subset of …

DGX agent

Big shoutout to @perplexity_ai for properly benchmarking document understanding in their Portable Computer release 🔥 They took a subset of our ParseBench benchmark (https://www.parsebench.ai/) and mea

model-releasesjerry-liu--x
26 Aug 2026
Model Releases

C3VDReg: A Benchmark for Local-to-Local Colonoscopic Registration toward Anatomical Localization

DGX agent

arXiv:2511.00260v2 Announce Type: replace Abstract: Anatomy-aware colonoscopic navigation requires localizing partial endoscopic observations on a stable 3D reference to support coverage assessment, r

model-releasesarxiv-cs-cv
26 Aug 2026
Model Releases

Calibration-Preserving Pruning: Compression as a Reliability Contract

DGX agent

arXiv:2608.23744v1 Announce Type: new Abstract: Split conformal prediction, not the pruning rule, supplies finite-sample marginal coverage once a pruned model is fixed independently of the conformal c

model-releasesarxiv-cs-lg
26 Aug 2026
Agents

Confident at the moment of action: belief miscalibration in LLM play under hidden information

DGX agent

arXiv:2608.24691v1 Announce Type: new Abstract: Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We te

agentsarxiv-cs-ai
26 Aug 2026
Research

Constrained Hyperparameter Optimization for Streaming Data

DGX agent

arXiv:2608.24712v1 Announce Type: cross Abstract: Optimization of hyperparameters is a critical factor to obtain optimal model performance. While existing research has predominantly concentrated on ba

researcharxiv-cs-ai
26 Aug 2026
Model Releases

Continual Visual Learning under Evolving Semantic Concept Shift

DGX agent

arXiv:2608.23903v1 Announce Type: new Abstract: Visual foundation models are commonly adapted under the assumption that the appearance of incoming data may change while the semantic meaning of the pre

model-releasesarxiv-cs-cv
26 Aug 2026
Local Ai

CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Semantic Segmentation

DGX agent

arXiv:2608.23746v1 Announce Type: new Abstract: State space models, especially Visual State Space Duality (VSSD), have emerged as efficient linear-time alternatives to Transformers for dense visual ta

local-aiarxiv-cs-cv
26 Aug 2026
Model Releases

Evaluating Multiple LLM Generations with Validated Task Coverage

DGX agent

arXiv:2608.24228v1 Announce Type: new Abstract: Many LLM applications are most useful when they provide several candidate outputs for comparison, validation, or combination. Predominant evaluation set

model-releasesarxiv-cs-ai
26 Aug 2026
Research

Giraffe: A Mapping Architecture from Hidden Text Representations to Visual Embeddings for Efficient Graphic Design

DGX agent

arXiv:2608.23970v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in understanding and interpreting mul- timedia content. However, their ability t

researcharxiv-cs-ai
26 Aug 2026
Model Releases

Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B…

DGX agent

Introducing GLM-5.3-Flash - Leading capabilities at a highly competitive price - Natively multimodal with a 1M-token context window - A 320B-A18B model released under the MIT License - Previously prev

model-releaseszhipu-ai--x
26 Aug 2026
Model Releases

Life-Bench: A Benchmark and Knowledge Graph Framework for Multimodal Personalization Beyond Concept Recognition

DGX agent

arXiv:2602.19001v2 Announce Type: replace Abstract: As large language models increasingly power personal assistants, users expect them to reason over multimodal life histories, from recognizing people

model-releasesarxiv-cs-cv
26 Aug 2026
Model Releases

Low-Rank Ternary Adaptation for Fine-Tuning Transformers

DGX agent

arXiv:2608.24469v1 Announce Type: new Abstract: Ternary transformers offer extreme memory and compute efficiency, but existing low-bit LoRA-based methods cannot directly fine-tune ternary weights. Cur

model-releasesarxiv-cs-cv
26 Aug 2026
Model Releases

MARS: Multi-Specialist LLM Relay System for Competitive Programming

DGX agent

arXiv:2608.23918v1 Announce Type: new Abstract: Large Language Models excel at code generation, yet competitive programming exposes a persistent failure mode: existing multi-agent pipelines distribute

model-releasesarxiv-cs-ai
26 Aug 2026
Model Releases

MatReplace: A Reference-Free, Conditioning-Aligned Benchmark for Material Replacement in Interior Scenes

DGX agent

arXiv:2608.24107v1 Announce Type: cross Abstract: Material replacement is a common interior-design operation: changing the material of a selected surface while preserving its geometry, surroundings, a

model-releasesarxiv-cs-ai
26 Aug 2026
Model Releases

[Megathread] Qwen3.8-Flash-Next - Release Day

DGX agent

Megathread for discussing the (impending) release of Qwen 3.8 Flash Next. Quants Fine-Tunes & Abliterations Chat Templates Inference Server Support & Configuration Experiences, Benchmarks & Model Comp

model-releasesr-localllama
26 Aug 2026
Research

MnemoDyn: Learning Resting State Dynamics from 40K FMRI sequences

DGX agent

arXiv:2608.23936v1 Announce Type: new Abstract: We present a dynamical-systems based model for resting-state functional magnetic resonance imaging (rs-fMRI), trained on a dataset of roughly 40K rs-fMR

researcharxiv-cs-lg
26 Aug 2026
Research

Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs

DGX agent

arXiv:2608.23897v1 Announce Type: cross Abstract: When a code generating language model fabricates a Python package name, an adversary who has pre-registered that name on PyPI can convert that halluci

researcharxiv-cs-ai
26 Aug 2026
Safety

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

DGX agent

arXiv:2608.24310v1 Announce Type: new Abstract: Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction,

safetyarxiv-cs-ai
26 Aug 2026
Model Releases

ORBITALIF: An Efficient Spiking Federated Learning Framework for Onboard Cloud Removal

DGX agent

arXiv:2608.24073v1 Announce Type: cross Abstract: Low-earth-orbit (LEO) satellites enable high-resolution, large-scale Earth observation for applications such as disaster monitoring and environmental

model-releasesarxiv-cs-ai
26 Aug 2026
Research

PonderPounce: A Pretrained MLLM as an Episode Context Engine for Robot Control

DGX agent

arXiv:2608.24115v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) can integrate long visual histories, reason under partial observability, and infer behavior from a few exampl

researcharxiv-cs-ai
26 Aug 2026
Research

Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems

DGX agent

arXiv:2608.23906v1 Announce Type: new Abstract: Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including Critical National Infrastructure (CNI), where har

researcharxiv-cs-ai
26 Aug 2026
Applications

Real-World Knowledge-Guided Change Data Synthesis for Remote Sensing

DGX agent

arXiv:2608.24263v1 Announce Type: new Abstract: Change data synthesis provides a cost-effective solution for expanding training data and improving the performance of change detection models. However,

applicationsarxiv-cs-ai
26 Aug 2026
Model Releases

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

DGX agent

arXiv:2608.24876v1 Announce Type: new Abstract: Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We in

model-releasesarxiv-cs-ai
26 Aug 2026
Model Releases

Robust Code RL via Faulty-Code-Driven Test case Synthesis and Dense Reward Shaping

DGX agent

arXiv:2608.24135v1 Announce Type: new Abstract: Reinforcement learning from verifiable rewards (RLVR) has emerged as a pivotal technique for enhancing the code generation capabilities of Large Languag

model-releasesarxiv-cs-ai
26 Aug 2026
← Previous
1…487488489490491…1371
Next →