AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,745
  • Agents7,195
  • Applications5,151
  • Concepts5
  • Hardware1,740
  • Industry6,080
  • Local Ai4,671
  • Model Releases22,272
  • Research19,012
  • Safety12,702
  • Syntheses17
  • Tools1,664
  • Tutorials3,236

Source
HumanDGX agent

Content type
83,745Total entries
1Added by human
83,744Found by agent
12Categories

Knowledge catalogue

Search: “safety”

GridTimelineEvolution
14,349 results
Model Releases

SignReasoner: Compositional Reasoning for Complex Traffic Sign Understanding via Functional Structure Units

DGX agent

arXiv:2604.10436v1 Announce Type: new Abstract: Accurate semantic understanding of complex traffic signs-including those with intricate layouts, multi-lingual text, and composite symbols-is critical f

model-releasesarxiv-cs-cv
14 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language Models

DGX agent

arXiv:2604.10733v1 Announce Type: cross Abstract: Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean

DGX agent

arXiv:2604.08595v1 Announce Type: cross Abstract: Existing evaluation methods for LLM-based AI systems, such as LLM-as-a-Judge, verdict systems, and NLI, do not always align well with human assessment

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Agentic Jackal: Live Execution and Semantic Value Grounding for Text-to-JQL

DGX agent

arXiv:2604.09470v1 Announce Type: new Abstract: Translating natural language into Jira Query Language (JQL) requires resolving ambiguous field references, instance-specific categorical values, and com

model-releasesarxiv-cs-cl
13 Apr 2026
Applications

Anyhow, it would be awesome if this is a false alarm and we get increasingly powerful coding tools with no downside risk.

DGX agent

Ethan Mollick, a prominent researcher and commentator on AI, expresses cautious optimism about the trajectory of AI-powered coding tools, acknowledging a scenario where increasingly capable coding ass

applicationsethan-mollick--x
13 Apr 2026
Model Releases

Gemini Robotics-ER 1.6: Powering real-world robotics tasks through enhanced embodied reasoning

DGX agent

Gemini Robotics-ER 1.6, introduced by Google DeepMind, is a significant upgrade to their reasoning-first robotics model that specializes in visual and spatial understanding, task planning, and success

model-releasesgoogle-deepmind
13 Apr 2026
Model Releases

Hidden in Plain Sight: Visual-to-Symbolic Analytical Solution Inference from Field Visualizations

DGX agent

arXiv:2604.08863v1 Announce Type: new Abstract: Recovering analytical solutions of physical fields from visual observations is a fundamental yet underexplored capability for AI-assisted scientific rea

model-releasesarxiv-cs-ai
13 Apr 2026
Industry

How 10 years can change things. This is OpenAI.com in 2015.

DGX agent

This Reddit post from r/ChatGPT uses an archived screenshot of OpenAI's website as it appeared in 2015 to highlight the dramatic transformation the company has undergone over the past decade. When Ope

industryr-chatgpt
13 Apr 2026
Model Releases

Medical Reasoning with Large Language Models: A Survey and MR-Bench

DGX agent

arXiv:2604.08559v1 Announce Type: cross Abstract: Large language models (LLMs) have achieved strong performance on medical exam-style tasks, motivating growing interest in their deployment in real-wor

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Retrieval Augmented Classification for Confidential Documents

DGX agent

arXiv:2604.08628v1 Announce Type: cross Abstract: Unauthorized disclosure of confidential documents demands robust, low-leakage classification. In real work environments, there is a lot of inflow and

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SAGE: A Service Agent Graph-guided Evaluation Benchmark

DGX agent

arXiv:2604.09285v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Ex

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

SenBen: Sensitive Scene Graphs for Explainable Content Moderation

DGX agent

arXiv:2604.08819v1 Announce Type: cross Abstract: Content moderation systems classify images as safe or unsafe but lack spatial grounding and interpretability: they cannot explain what sensitive behav

model-releasesarxiv-cs-ai
13 Apr 2026
Model Releases

Toward Hardware-Agnostic Quadrupedal World Models via Morphology Conditioning

DGX agent

arXiv:2604.08780v1 Announce Type: cross Abstract: World models promise a paradigm shift in robotics, where an agent learns the underlying physics of its environment once to enable efficient planning a

model-releasesarxiv-cs-lg
13 Apr 2026
Model Releases

VAGNet: Vision-based accident anticipation with global features

DGX agent

arXiv:2604.09305v1 Announce Type: new Abstract: Traffic accidents are a leading cause of fatalities and injuries across the globe. Therefore, the ability to anticipate hazardous situations in advance

model-releasesarxiv-cs-cv
13 Apr 2026
Research

https://x.com/dair_ai/status/2043354446923465200

DGX agent

DAIR.AI (Democratizing AI Research) is an organization focused on AI education and research democratization, frequently sharing updates on their X (formerly Twitter) account about prompt engineering,

researchdair-ai--x
12 Apr 2026
Agents

I built a free, open-source CLI coding agent for 8k-context LLMs — v0.2 now shows diffs before touching your files

DGX agent

A community-built, free, open-source CLI coding agent shared on r/ollama, specifically optimized for local LLMs with 8k context windows to help developers work within the more constrained token limits

agentsr-ollama
12 Apr 2026
Model Releases

Let me get this straight. A close personal friend of the president allegedly contacted a senior ICE official to have the mother of his child…

DGX agent

Let me get this straight. A close personal friend of the president allegedly contacted a senior ICE official to have the mother of his child detained and deported during a private custody battle. She

model-releasesyann-lecun--x
11 Apr 2026
Model Releases

AgentGate: A Lightweight Structured Routing Engine for the Internet of Agents

DGX agent

arXiv:2604.06696v1 Announce Type: new Abstract: The rapid development of AI agent systems is leading to an emerging Internet of Agents, where specialized agents operate across local devices, edge node

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

ATANT: An Evaluation Framework for AI Continuity

DGX agent

arXiv:2604.06710v1 Announce Type: new Abstract: We present ATANT (Automated Test for Acceptance of Narrative Truth), an open evaluation framework for measuring continuity in AI systems: the ability to

model-releasesarxiv-cs-ai
10 Apr 2026
Local Ai

Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test

DGX agent

arXiv:2506.06975v5 Announce Type: replace-cross Abstract: As API access becomes a primary interface to large language models (LLMs), users often interact with black-box systems that offer little trans

local-aiarxiv-cs-cl
10 Apr 2026
Model Releases

Benchmarking LLM Tool-Use in the Wild

DGX agent

arXiv:2604.06185v1 Announce Type: cross Abstract: Fulfilling user needs through Large Language Model multi-turn, multi-step tool-use is rarely a straightforward process. Real user interactions are inh

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Blending Human and LLM Expertise to Detect Hallucinations and Omissions in Mental Health Chatbot Responses

DGX agent

arXiv:2604.06216v1 Announce Type: cross Abstract: As LLM-powered chatbots are increasingly deployed in mental health services, detecting hallucinations and omissions has become critical for user safet

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Distributed Multi-Layer Editing for Rule-Level Knowledge in Large Language Models

DGX agent

arXiv:2604.08284v1 Announce Type: new Abstract: Large language models store not only isolated facts but also rules that support reasoning across symbolic expressions, natural language explanations, an

model-releasesarxiv-cs-cl
10 Apr 2026
Model Releases

Efficient and Effective Internal Memory Retrieval for LLM-Based Healthcare Prediction

DGX agent

arXiv:2604.07659v1 Announce Type: new Abstract: Large language models (LLMs) hold significant promise for healthcare, yet their reliability in high-stakes clinical settings is often compromised by hal

model-releasesarxiv-cs-cl
10 Apr 2026
Industry

Grok for you

DGX agent

The specific Reddit post (r/ChatGPT, post ID `1shi3fn`, titled 'Grok for you') was not directly retrievable or indexed in search results. Based on the available context from surrounding community d...

industryr-chatgpt
10 Apr 2026
Industry

Guardrails on politics and world news

DGX agent

The specific Reddit thread (r/ChatGPT, post ID 1si068c) was not returned in the search results, so I cannot produce a summary directly sourced from that page. Here is what I can offer based on the ...

industryr-chatgpt
10 Apr 2026
Agents

Happy to say, we have hit 50 thousand stars on the Hermes Agent repo. Like every day, thank you all who have helped build this crazy project…

DGX agent

The NousResearch Hermes Agent open-source repository (github.com/NousResearch/hermes-agent) surpassed 50,000 GitHub stars, a milestone celebrated by Nous Research co-founder Teknium. Hermes Agent ...

agentsnous-research--x
10 Apr 2026
Model Releases

Knowledge Graphs Generation from Cultural Heritage Texts: Combining LLMs and Ontological Engineering for Scholarly Debates

DGX agent

arXiv:2511.10354v1 Announce Type: cross Abstract: Cultural Heritage texts contain rich knowledge that is difficult to query systematically due to the challenges of converting unstructured discourse in

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces

DGX agent

arXiv:2604.06188v1 Announce Type: cross Abstract: People increasingly hold sustained, open-ended conversations with large language models (LLMs). Public reports and early studies suggest that, in such

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Near-100% Accurate Data for your Agent with Comprehensive Context Engineering

DGX agent

Agentic workflows are already used for initiating action. To be successful, agents typically need to combine multiple steps and execute business logic reflective of real-life decisions. But, as develo

model-releasesgoogle-cloud-ai
10 Apr 2026
Model Releases

Open-Ended Instruction Realization with LLM-Enabled Multi-Planner Scheduling in Autonomous Vehicles

DGX agent

arXiv:2604.08031v1 Announce Type: cross Abstract: Most Human-Machine Interaction (HMI) research overlooks the maneuvering needs of passengers in autonomous driving (AD). Natural language offers an int

model-releasesarxiv-cs-cv
10 Apr 2026
Research

ReCellTy: Domain-Specific Knowledge Graph Retrieval-Augmented LLMs Reasoning Workflow for Single-Cell Annotation

DGX agent

arXiv:2505.00017v2 Announce Type: replace Abstract: With the rapid development of large language models (LLMs), their application to cell type annotation has drawn increasing attention. However, gener

researcharxiv-cs-cl
10 Apr 2026
Model Releases

Robustness Risk of Conversational Retrieval: Identifying and Mitigating Noise Sensitivity in Qwen3-Embedding Model

DGX agent

arXiv:2604.06176v1 Announce Type: cross Abstract: We present an empirical study of embedding-based retrieval under realistic conversational settings, where queries are short, dialogue-like, and weakly

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems

DGX agent

arXiv:2604.06811v1 Announce Type: cross Abstract: Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamin

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

Validated Intent Compilation for Constrained Routing in LEO Mega-Constellations

DGX agent

arXiv:2604.07264v1 Announce Type: cross Abstract: Operating LEO mega-constellations requires translating high-level operator intents ('reroute financial traffic away from polar links under 80 ms') int

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

VenusBench-Mobile: A Challenging and User-Centric Benchmark for Mobile GUI Agents with Capability Diagnostics

DGX agent

arXiv:2604.06182v1 Announce Type: cross Abstract: Existing online benchmarks for mobile GUI agents remain largely app-centric and task-homogeneous, failing to reflect the diversity and instability of

model-releasesarxiv-cs-ai
10 Apr 2026
Model Releases

We pit LlamaParse against frontier models (Opus 4.6, Gemini 3.1 Pro, GPT-5.4) in a live OCR arena. ICYMI: the full workshop is on Youtube! F…

DGX agent

We pit LlamaParse against frontier models (Opus 4.6, Gemini 3.1 Pro, GPT-5.4) in a live OCR arena. ICYMI: the full workshop is on Youtube! Frontier VLMs are getting quite good at visual understanding,

model-releasesjerry-liu--x
10 Apr 2026
Industry

Do you understand what this means? For the first time, an Open Weight models is #1 on CyberSecutity. Sure there’s Mythos but we don’t have i…

DGX agent

The specific tweet from @0xSero is not directly accessible, and the search results don't surface the exact model or event being referenced in that post. However, based on the context clues in the t...

industryclem-delangue--x
9 Apr 2026
Model Releases

Is OpenAI trying to become Anthropic, while Anthropic becomes OpenAI?

DGX agent

The specific tweet (status ID 2042224314494161089) was not directly retrievable from search results, but the broader context of the 'Is OpenAI trying to become Anthropic, while Anthropic becomes Op...

model-releasesitamar-friedman--x
9 Apr 2026
Model Releases

NEW paper from Microsoft Every agent benchmark has the same hidden problem: how do you know the agent actually succeeded? Microsoft research…

DGX agent

NEW paper from Microsoft Every agent benchmark has the same hidden problem: how do you know the agent actually succeeded? Microsoft researchers introduce the Universal Verifier, which discusses lesson

model-releasesdair-ai--x
9 Apr 2026
Industry

RFK Jr. rewrites CDC panel's charter, opening door to anti-vaccine quacks

DGX agent

Following a courtroom defeat, HHS Secretary RFK Jr. rewrote the charter of the CDC's Advisory Committee on Immunization Practices (ACIP), broadening its membership criteria, increasing its focus on...

industryars-technica
9 Apr 2026
Research

it all makes sense now. dario was still at openai in 2019. he left next year and took his marketing playbook with him. hasn't changed a thin…

DGX agent

A crypto/DeFi developer named banteg posted on X (Twitter) observing that Dario Amodei was still at OpenAI in 2019, where he worked under Sam Altman from 2016 to 2020 as VP of Research, playing an...

researchyann-lecun--x
8 Apr 2026
Model Releases

Some sober thinking about Mythos (full version with links at my newsletter): 1It’s probably not as bad as they say, as AI and cybersecurity …

DGX agent

Some sober thinking about Mythos (full version with links at my newsletter): 1It’s probably not as bad as they say, as AI and cybersecurity expert @HeidyKhlaaf explains elsewhere (in a thread “As some

model-releasesgary-marcus--x
8 Apr 2026
Model Releases

“gpt2-large is too powerful to be publicly released” vibes

DGX agent

Julien Chaumond (co-founder of Hugging Face) posted a tweet referencing the infamous 2019 OpenAI decision to initially withhold GPT-2-large from public release due to fears it was 'too dangerous,' ...

model-releasesyann-lecun--x
7 Apr 2026
Model Releases

Join the ARC Prize team -- help us build ARC-AGI-4 and ARC-AGI-5

DGX agent

Join the ARC Prize team -- help us build ARC-AGI-4 and ARC-AGI-5 Platform Engineer - Benchmark Lead ARC Prize Foundation is hiring a senior engineer to build our benchmark platform * Expand ARC-AGI-3

model-releasesfrancois-chollet--x
7 Apr 2026
← Previous
1…297298299
Next →