AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,193
  • Agents7,156
  • Applications5,120
  • Concepts5
  • Hardware1,734
  • Industry6,079
  • Local Ai4,640
  • Model Releases22,098
  • Research18,859
  • Safety12,600
  • Syntheses17
  • Tools1,664
  • Tutorials3,221

Source
HumanDGX agent

Content type
AllBlog
83,193Total entries
1Added by human
83,192Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,237 results
Model Releases

SymStep: Symbolic Step Verification for Logical Reasoning

DGX agent

arXiv:2607.23055v1 Announce Type: new Abstract: Chain-of-thought (CoT) prompting can fail severely on constraint-dense logical reasoning tasks, where unverified errors accumulate silently across steps

model-releasesarxiv-cs-ai
28 Jul 2026
X Post
Paper
YouTube
Reddit
GitHub
Clear filters
Model Releases

The Half-Lives of Generative-AI Evidence: A 40-Record Audit, a Claim-Currency Framework, and a Reflexive Case of Frontier-Model-Assisted Research

DGX agent

arXiv:2607.24032v1 Announce Type: new Abstract: Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked

model-releasesarxiv-cs-ai
28 Jul 2026
Model Releases

Toward Automated Detection of Documentation Inconsistencies in Electronic Health Records

DGX agent

arXiv:2607.22954v1 Announce Type: new Abstract: Objective: To characterize the kinds of internal documentation inconsistencies a general-domain large language model (LLM) can surface from real-world d

model-releasesarxiv-cs-cl
28 Jul 2026
Model Releases

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

DGX agent

arXiv:2607.22727v1 Announce Type: new Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can

model-releasesarxiv-cs-cv
28 Jul 2026
Model Releases

UP-Fuse: Uncertainty-guided LiDAR-Camera Fusion for 3D Panoptic Segmentation

DGX agent

arXiv:2602.19349v2 Announce Type: replace-cross Abstract: LiDAR-camera fusion enhances 3D panoptic segmentation by leveraging camera images to complement sparse LiDAR scans, but it also introduces a c

model-releasesarxiv-cs-ai
28 Jul 2026
Local Ai

VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

DGX agent

arXiv:2607.23006v1 Announce Type: cross Abstract: Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the suppo

local-aiarxiv-cs-ai
28 Jul 2026
Model Releases

Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study

DGX agent

arXiv:2607.21988v1 Announce Type: new Abstract: Self-harm content is particularly challenging to detect using NLP techniques, and is also a high-stakes task which requires the highest accuracy to enab

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

DBA-Bench: A Production-Fidelity Benchmark for LLM-Based Database Operations Agents

DGX agent

arXiv:2607.22165v1 Announce Type: cross Abstract: LLM-based database agents show promise, but differing task scopes, testbeds, and metrics hinder comparison. We identify four gaps between evaluation a

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

Microsoft introduces MAI-Cyber-1-Flash, an AI model trained for cybersecurity, and launches Perception, an agentic security system to patch vulnerabilities (New York Times)

DGX agent

New York Times: Microsoft introduces MAI-Cyber-1-Flash, an AI model trained for cybersecurity, and launches Perception, an agentic security system to patch vulnerabilities — As some executives fret ov

model-releasestechmeme
27 Jul 2026
Model Releases

Modernizing the skies: NOAA and Google Cloud collaborate to advance weather forecasting

DGX agent

The National Oceanic and Atmospheric Administration (NOAA) is embarking on a transformative journey to redefine how we understand and predict patterns in the Earth’s atmosphere that affect the weather

model-releasesgoogle-cloud-ai
27 Jul 2026
Model Releases

Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science

DGX agent

arXiv:2607.22513v1 Announce Type: cross Abstract: Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor

model-releasesarxiv-cs-cl
27 Jul 2026
Model Releases

The 3D Mirage: Probing and Taming 3D Hallucinations

DGX agent

arXiv:2512.15423v2 Announce Type: replace Abstract: Monocular depth foundation models achieve remarkable generalization by learning large-scale semantic priors, but this creates a critical vulnerabili

model-releasesarxiv-cs-cv
27 Jul 2026
Model Releases

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

DGX agent

arXiv:2607.21635v1 Announce Type: new Abstract: Personal agents maintain memories, learned skills, tool configurations, and policy state that evolve with each user. Existing agent benchmarks often eva

model-releasesarxiv-cs-lg
27 Jul 2026
Model Releases

23 Gemma4-E4B models compared with abliterlitics: the most downloaded one is also the most broken

DGX agent

This is our biggest comparison yet. We've taken 23 Gemma 4 E4B models from huggingface and ran them through the abliterlitics gauntlet. We also have a new abliterlitics discord, feel free to jump on a

model-releasesr-localllama
26 Jul 2026
Model Releases

Nvidia, other tech giants caution against open-source AI ban in open letter

DGX agent

A group of tech firms has released an open letter that calls on policymakers not to ban open-source artificial intelligence models. The development follows a report that some Trump administration offi

model-releasessiliconangle
25 Jul 2026
Model Releases

Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact hav…

DGX agent

Very happy to support this on behalf of Google. We have long benefited from open source, are big contributors to open source and in fact have consistently made open weights models with Gemma available

model-releasesclem-delangue--x
25 Jul 2026
Model Releases

Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants

DGX agent

arXiv:2607.20488v1 Announce Type: new Abstract: Multi-agent LLM frameworks typically fix their team topology at boot time. When an individual agent becomes overloaded at runtime, for example by mixing

model-releasesarxiv-cs-ai
24 Jul 2026
Local Ai

Concept Concentration for Faithful Representation Intervention

DGX agent

arXiv:2505.18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to e

local-aiarxiv-cs-lg
24 Jul 2026
Local Ai

Conflict Resolution under Degraded Surveillance in Air Corridors Using Multi-Agent Reinforcement Learning

DGX agent

arXiv:2607.20547v1 Announce Type: new Abstract: Safe Advanced Air Mobility operations require aircraft to maintain separation when surveillance information is noisy, delayed, incomplete, or temporaril

local-aiarxiv-cs-lg
24 Jul 2026
Model Releases

Geometric Configurations of Perturbed Jailbreak Prompts

DGX agent

arXiv:2607.20581v1 Announce Type: cross Abstract: Perturbation techniques that turn unsuccessful jailbreak prompts into successful ones are continuously evolving, constituting a major security threat

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

GuardianAgentBench: Where Agents Fail and How to Guard Them

DGX agent

arXiv:2607.20982v1 Announce Type: new Abstract: As large language model agents increasingly operate autonomously with access to tools and external environments, ensuring their safe and reliable behavi

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

ImplicitBBQ: Benchmarking Implicit Bias in Large Language Models through Characteristic Based Cues

DGX agent

arXiv:2604.01925v2 Announce Type: replace-cross Abstract: Large Language Models increasingly suppress biased outputs when demographic identity is stated explicitly, yet may still exhibit implicit bias

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests

DGX agent

arXiv:2607.20759v1 Announce Type: cross Abstract: AI coding agents powered by LLMs are increasingly integrated into real-world software development, where they generate, edit, and execute code with au

model-releasesarxiv-cs-ai
24 Jul 2026
Local Ai

Open models matter. Ollama works hard with the model creators, hardware partners, and most importantly developers building software leveragi…

DGX agent

Open models matter. Ollama works hard with the model creators, hardware partners, and most importantly developers building software leveraging various open models for their own use cases. For my first

local-aiollama--x
24 Jul 2026
Local Ai

Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your p…

DGX agent

Open weights = freedom. You can run them on your own hardware. No vendor can pull the plug. No API can deprecate you. No company logs your private data. That's sovereignty. Closed models hand one comp

local-aifireworks-ai--x
24 Jul 2026
Model Releases

Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling

DGX agent

arXiv:2607.20791v1 Announce Type: new Abstract: High-temperature sampling is one of the primary mechanisms for increasing diversity in LLMs. Recent advances in truncation-based sampling techniques hav

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation

DGX agent

arXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction.

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

SkillCorpus: Consolidating and Evaluating the Open Skill Ecosystem for Real-World LLM Agents

DGX agent

arXiv:2607.15557v4 Announce Type: replace Abstract: Agent skills, SKILL files that package reusable procedural knowledge for an LLM agent, are a popular mechanism for extending agent capabilities. Pub

model-releasesarxiv-cs-cl
24 Jul 2026
Model Releases

StabilityBench: Benchmarking Instability in LLMs

DGX agent

arXiv:2607.20558v1 Announce Type: cross Abstract: AI Assistants are increasingly deployed in high-stakes settings, such as healthcare or government services. Yet their real-world behavior remains poor

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning

DGX agent

arXiv:2607.09328v2 Announce Type: replace-cross Abstract: Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across dis

model-releasesarxiv-cs-ai
24 Jul 2026
Model Releases

ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems

DGX agent

arXiv:2607.19430v1 Announce Type: cross Abstract: Multi-agent LLM applications chain a planner, worker agents, a verifier, and a synthesizer, and every hop between agents is an unmonitored channel thr

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Emergent Autonomous Drifting for Collision Avoidance in Real-World Winter Driving Scenarios

DGX agent

arXiv:2607.19484v1 Announce Type: new Abstract: Real-world collision avoidance is a core motivation for studying the dynamics and control of high sideslip drifting in vehicles, yet the practical benef

model-releasesarxiv-cs-ro
23 Jul 2026
Model Releases

FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance

DGX agent

arXiv:2607.19409v1 Announce Type: new Abstract: Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasize measuring

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models

DGX agent

arXiv:2607.19424v1 Announce Type: cross Abstract: The assessment of jailbreak attacks against large language models currently suffers from inconsistent evaluation criteria and methods, leading to unre

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

The Blueprint: How Voicify makes AI-enabled ordering a delight for customers

DGX agent

Welcome to The Blueprint, a new feature where we highlight how Google Cloud customers are tackling unique and common challenges across industries using the latest AI and cloud technologies. We hope to

model-releasesgoogle-cloud-ai
23 Jul 2026
Model Releases

Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

DGX agent

arXiv:2607.20379v1 Announce Type: new Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be reg

model-releasesarxiv-cs-ai
23 Jul 2026
Model Releases

Unified Prediction and Planning via Conflict-Aware Disjoint Parameter Training

DGX agent

arXiv:2607.19971v1 Announce Type: new Abstract: Accurate motion prediction of surrounding agents and safe motion planning are two closely coupled key tasks for social robot navigation in crowded envir

model-releasesarxiv-cs-ro
23 Jul 2026
Model Releases

Instead of panicking about the Hugging Face attack, people need to start questioning OpenAI's insecure sandboxes.

DGX agent

One thing I noticed in American politics, whenever the government wants to push unpopular actions or laws, they often introduce fear to convince the public to support them. This is actually how i view

model-releasesr-localllama
22 Jul 2026
Model Releases

OpenAI’s zero-day exploit hack of HuggingFace *should* be a wake up call. Although there lots of caveats around what happened, we are just g…

DGX agent

OpenAI’s zero-day exploit hack of HuggingFace *should* be a wake up call. Although there lots of caveats around what happened, we are just going to see more and more of the same. We have no guarantees

model-releasesgary-marcus--x
22 Jul 2026
Model Releases

Stuck scaling a Next.js app on M3 Pro (36GB) using local Qwen 3.6 + VS Code Copilot. Should I switch extensions or go paid?

DGX agent

Hey everyone, I’m a Full-Stack Developer with 6+ years of experience. I’m relatively new to AI-assisted development workflows and want to build a production-ready, enterprise-level Next.js web applica

model-releasesr-ollama
22 Jul 2026
Local Ai

I just wanted a small WebUI with an admin panel… it escalated into a full open-source agent framework runs fully local with Ollama

DGX agent

Let me try to explain this clearly, simply, and neatly. Originally, I just wanted to build a small WebUI adapter with an admin panel, but things escalated over the last few months. At first, I faced t

local-air-ollama
21 Jul 2026
Industry

Last Week in AI #250 - Mythos Mess, GPT 5.6-Sol, GLM 5.2

DGX agent

Last Week in AI #250 is the podcast’s 250th episode, summarizing recent frontier‑AI policy developments. It reports that the U.S. government has granted Anthropic permission to release Mythos‑5 to a l

industrylast-week-in-ai
21 Jul 2026
Model Releases

LWiAI Podcast #248 - Opus 4.8, MAI, Anthropic IPO, Minimax-M3

DGX agent

LWiAI Podcast #248 (June 12, 2026) reviews major AI developments, noting Anthropic’s release of Claude Fable 5, a safeguarded variant of Mythos 5, which shows benchmark improvements but raises concern

model-releaseslast-week-in-ai
21 Jul 2026
Model Releases

LWiAI Podcast #252 - GPT 5.6, Grok 4.5, Nemotron-Labs-Diffusion, AI 2040

DGX agent

LWiAI Podcast #252 (July 11, 2026) reviewed major AI releases: OpenAI unveiled GPT‑5.6 and relaunched its agentic coding product as ChatGPT Work, amid disputes over U.S. governmental oversight and jai

model-releaseslast-week-in-ai
21 Jul 2026
Syntheses

Wiki Lint Report — 2026-07-19

DGX agent

Automated lint: 20 errors, 8743 warnings, 3 info

linthealth-checkautomated
19 Jul 2026
Model Releases

A Self-Evolving Agent for Longitudinal Personal Health Management

DGX agent

arXiv:2607.13940v1 Announce Type: new Abstract: Personal health management unfolds over repeated encounters, yet most health AI systems treat each request in isolation. We developed HealthClaw, an ope

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

CAVA: Canonical Action Verification and Attestation for Runtime Governance of Agentic AI Systems

DGX agent

arXiv:2607.13716v1 Announce Type: new Abstract: Agentic AI systems increasingly act through heterogeneous runtimes: local coding hooks, SDK tools, browser automation, managed-agent traces, API gateway

model-releasesarxiv-cs-ai
16 Jul 2026
Model Releases

Efficient Text-to-Audio Generation via Pruning

DGX agent

arXiv:2607.13330v1 Announce Type: cross Abstract: Diffusion-based text-to-audio generative models such as AudioLDM achieve high perceptual quality and strong semantic consistency; however, their pract

model-releasesarxiv-cs-ai
16 Jul 2026
← Previous
1…283284285286287…297
Next →