AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “tools”

GridTimelineEvolution
10,103 results
12 Aug 2026

From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents

Model ReleasesDGX agent

arXiv:2608.10502v1 Announce Type: new Abstract: Persistent memory lets language-model agents reuse information across sessions, but it also makes errors durable: a poisoned, stale, or misattributed re

Idea for a deepseek-v4-flash-0731 backed automated research workflow to be leveraged via qwen3.6/3.8 27b for difficult tasks that require highly technical, not easy to find information.

Model ReleasesDGX agent

Sometimes you have tasks that are outside of your expertise and the idea is this workflow automation could be leveraged to manage to have local AI figure it out using research from his workflow gather

On The Statistical Limits of Self-Improving Agents

SafetyDGX agent
Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

arXiv:2510.04399v3 Announce Type: replace Abstract: We develop a learning-theoretic framework for analyzing self-improving agents by decomposing self-modification into five axes. Within this framework

RLMOpt: Adaptive Prompt Optimization via Recursive Language Models

Model ReleasesDGX agent

arXiv:2608.10471v1 Announce Type: new Abstract: Prompt optimizers automate the search for prompts that improve language-model performance, but existing methods rely on a predefined optimization proced

SkillZip: Evaluation-Free Skill Compression for Self-Evolving Agents by Discovering Reusable Structure

ResearchDGX agent

arXiv:2608.11079v1 Announce Type: new Abstract: Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated

The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse

AgentsDGX agent

arXiv:2608.10186v1 Announce Type: cross Abstract: LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests l

The GenAI Catch-22: Use of Generative Artificial Intelligence in Norwegian Newsrooms During the 2025 Parliamentary Election

ApplicationsDGX agent

arXiv:2608.10773v1 Announce Type: cross Abstract: The increasing use of Generative Artificial Intelligence (GenAI) in journalism raises concerns about possible detrimental effects both on journalism a

When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning

Model ReleasesDGX agent

arXiv:2608.09942v1 Announce Type: cross Abstract: It is widely assumed that chain-of-thought (CoT) prompting universally improves LLM reasoning. We investigate this through the conceptual framework of

11 Aug 2026

Accelerate PostgreSQL migrations using Gemini in Database Migration Service

Model ReleasesDGX agent

Imagine this scenario: Your team decides to migrate a core application from an existing commercial database like Oracle or SQL Server to open source PostgreSQL or a fully managed service such as Alloy

APEX-VW: A Document-Level English-Spanish Post-Editing Dataset in the Healthcare Domain

Model ReleasesDGX agent

arXiv:2608.08059v1 Announce Type: new Abstract: Post-Editing (PE) of Machine Translation (MT) output often involves repeating the same lexical and terminological corrections across many segments, espe

BibTeX Citation Errors in Scientific Publishing Agents: Evaluation and Mitigation

Model ReleasesDGX agent

arXiv:2604.03159v2 Announce Type: replace-cross Abstract: Large language models with web search are increasingly used in scientific publishing agents, yet they produce BibTeX entries with pervasive fi

Can Gemma and Qwen models catch hallucinations by looking at their own logprobs?

Model ReleasesDGX agent

Hi! I'm really obsessed with LLM hallucinations for the last 6 days 😭 I started by designing system prompts to attack hallucinations but failed, obviously. Now I tried reading logprobs and... I think

DocAtlas: Long-Document Understanding as Mutable-State Interaction

Model ReleasesDGX agent

arXiv:2608.07527v1 Announce Type: cross Abstract: Long-document understanding requires models to find and combine evidence across many pages, layouts, tables, figures, and charts. Existing retrieval-a

Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books

ApplicationsDGX agent

arXiv:2608.08349v1 Announce Type: cross Abstract: Audio dramas weave dialogue, sound effects, and music into immersive stories. Creators often adapt books into audio dramas, but this process remains l

From Prompt to Harness: Coderlet from Scratch

AgentsDGX agent

arXiv:2608.09480v1 Announce Type: new Abstract: A model alone does not determine how a programming agent acts. What the model sees, how actions enter the environment, how feedback returns, and how one

HandSplatter: Automated Digital Goniometry from Neural Rendering

ResearchDGX agent

arXiv:2608.09735v1 Announce Type: new Abstract: Hand and finger disorders are leading contributors to musculoskeletal disability, creating a clinical need for precise methods to quantify joint motion.

Looker’s semantic layer governs Gemini Enterprise data for user trust

Model ReleasesDGX agent

For organizations deploying AI agents at scale, there’s often a critical divide between structured and unstructured data. While large language models (LLMs) excel at parsing text documents, emails, an

Snowflake moves enterprise AI beyond fragmented data pipelines

ApplicationsDGX agent

Data interoperability is quickly becoming a practical requirement for companies trying to move artificial intelligence into production. Picking the right model or adding computing capacity is only par

Thinking Is Not Telling: Information Disclosure in User-Service LLM Agents

AgentsDGX agent

arXiv:2602.07796v2 Announce Type: replace Abstract: User-engaged LLM agents increasingly operate in service scenarios where task success depends on coordination between the agent, the user, and a stat

Towards Researcher Agents for Knowledge-Graph Question Answering

Model ReleasesDGX agent

arXiv:2608.07700v1 Announce Type: new Abstract: Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, g

Unsure but Certain: Uncovering the Representation-Confidence Gap in Diffusion Language Models

ResearchDGX agent

arXiv:2608.08791v1 Announce Type: new Abstract: Diffusion language models use broad context to create text, suggesting they might handle input noise better than standard models. Testing reveals this i

Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning

Local AiDGX agent

arXiv:2608.07593v1 Announce Type: cross Abstract: Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Exi

10 Aug 2026

ADIAS: Automated Design of Interactive Agentic Systems

AgentsDGX agent

arXiv:2608.06410v1 Announce Type: new Abstract: Automated agent design improves agent harnesses through iterative revision, evaluation, and feedback summarization. Existing methods are largely candida

Best open-source harness like Claude Code?

Model ReleasesDGX agent

Avid claude code user here looking to do equivalent things with local models. Just want to plug in something like Qwen and have the interface be 1:1 with claude code. Any suggestion? submitted by /u/N

Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence

Model ReleasesDGX agent

arXiv:2608.06756v1 Announce Type: new Abstract: Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes

Comparing how Cline, Kilo, and Qwen Code handle long-task context/state (and why context loops keep happening)

Model ReleasesDGX agent

I've been comparing Cline / Kilo / Qwen Code lately since they all handle long-task state differently. Cline: has Focus Chain, a markdown file kept outside the conversation that gets reinjected on a c

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

Model ReleasesDGX agent

arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t

Homebot: A Personal AI Agent for Conversational Home Assistance and Automation

Local AiDGX agent

arXiv:2608.02254v2 Announce Type: replace Abstract: exttt{Homebot} is a locally deployable AI agent for conversational household assistance and automation. It accepts voice and instant-messaging reque

How WPP operationalizes platform and data engineering for AI marketing

SafetyDGX agent

Between chaotic levels of market fragmentation and economic volatility, marketing and communications agencies can no longer rely on the human intuition they’ve traditionally used to win clients and op

Not All Problems Are Best Modeled as MILP: A DSL-Centric Framework for Flexible and Accurate Optimization Modeling

Model ReleasesDGX agent

arXiv:2608.07040v1 Announce Type: new Abstract: Solving combinatorial optimization problems (COPs) requires not only efficient algorithms but also carefully crafted formulations. While recent works ha

Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents

Model ReleasesDGX agent

arXiv:2608.06968v1 Announce Type: cross Abstract: Language-model agents act on state encodings of their environment, yet these are treated as interchangeable interfaces. Using pretrained language mode

The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML Workflows

SafetyDGX agent

arXiv:2608.06714v1 Announce Type: new Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits

There is an asymmetry in most agentic workflows that does not get talked about much: humans have many ways to talk to agents, and almost no …

Model ReleasesDGX agent

There is an asymmetry in most agentic workflows that does not get talked about much: humans have many ways to talk to agents, and almost no standardized way for agents to talk back to humans. You can

When Do LLMs Admit Their Mistakes? Understanding The Role Of Model Belief In Retraction

TutorialsDGX agent

arXiv:2505.16170v4 Announce Type: replace Abstract: We study the internal mechanisms that govern when LLMs choose to retract wrong answers, i.e., spontaneously and immediately acknowledge errors in th

8 Aug 2026

and we agree with @GaryMarcus. Training and running frontier class AI on CPUs, at a fractional cost, saving the planet, while building human…

SafetyDGX agent

and we agree with @GaryMarcus. Training and running frontier class AI on CPUs, at a fractional cost, saving the planet, while building human aligned AI is the 2nd innings of AI race. CPUs and the rise

Imagine image 2.0, non-agentic yet, more to come in a week or two 💙

AgentsDGX agent

Imagine image 2.0, non-agentic yet, more to come in a week or two 💙 Announcing Imagine Image 2.0, our next generation image model with precision editing, crisp text rendering, improved factuality, and

7 Aug 2026

CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation Prediction

Model ReleasesDGX agent

arXiv:2608.05359v1 Announce Type: new Abstract: CASCADE is an agentic framework that predicts downstream transcriptional effects of gene perturbation from precomputed ARACNe regulatory networks, expos

ECHO: A Locally-Deployable Agentic Health Assistant with Temporal Memory, Safety Guardrails, and Speech Assessment

Model ReleasesDGX agent

arXiv:2608.06110v1 Announce Type: new Abstract: This paper presents ECHO (Enhanced Care & Health Observer), a locally-deployable conversational health assistant for long-term chronic care management.

good grok

Model ReleasesDGX agent

good grok Best match for this hierarchical hands-free setup: - Runtime: ActiveGraph (event-sourced log as source of truth) or LangGraph for supervisor/manager graphs - Roles as skills: Claude Agent SD

Menlo Security targets real-time AI agent security with MARS platform

SafetyDGX agent

As AI agents gain access to enterprise systems and sensitive data, security teams need greater visibility into their actions. Real-time monitoring and policy enforcement are becoming essential as a pa

6 Aug 2026

Agents want to collaborate. So we’re putting that instinct to work on improving open-weight LLMs at formal math. Our latest agent collab tac…

AgentsDGX agent

Agents want to collaborate. So we’re putting that instinct to work on improving open-weight LLMs at formal math. Our latest agent collab tackles the @SAIRfoundation challenge of building a cheat sheet

Among proponents of neurosymbolic architectures, there had been some debate over the years about whether the outer level would be symbolic (…

SafetyDGX agent

Among proponents of neurosymbolic architectures, there had been some debate over the years about whether the outer level would be symbolic (i.e. a harness that calls neural models) or whether the oute

i just spent weeks rewriting my webUI from scratch, getting rid of all AI slop within the codebase and switching it over to a proper lightweight framework (alpine.js). i am now comfortable suggesting it as an alternative to openwebUI, librechat and the like! it is made for local models

Local AiDGX agent

[Fully open source under GPL3, made from the ground up for use with local models, no subscriptions, no corporate backing] When i first started this, it was meant to be a fully lightweight, extremely m

Monsoon Mayhem to Market Waves: Forecasting Fisheries Resilience in Sri Lanka

ApplicationsDGX agent

arXiv:2608.04023v1 Announce Type: cross Abstract: Sri Lanka's fisheries sector is important for jobs and food supply. Between 2019 and 2025, it faced several major problems at the same time, and how t

ReGround: Restoring Visual Grounding in Multi-Step Reasoning through Self-Diagnosis and Visual Re-Examination

Model ReleasesDGX agent

arXiv:2608.04385v1 Announce Type: new Abstract: Vision-Language Models (VLMs) often lose visual grounding during multi-step reasoning: as reasoning chains grow longer, later inference steps rely incre

Sources: Canva slashed revenue growth forecast as heavy use of new AI features drove up costs and slowed their rollout, while more Canva users turned to ChatGPT (The Information)

IndustryDGX agent

The Information: Sources: Canva slashed revenue growth forecast as heavy use of new AI features drove up costs and slowed their rollout, while more Canva users turned to ChatGPT — Executives at Canva,

5 Aug 2026

A Survey on Design Methodologies for Accelerating Deep Learning on Heterogeneous Architectures

ResearchDGX agent

arXiv:2311.17815v3 Announce Type: replace-cross Abstract: Given their increasing size and complexity, the need for efficient execution of deep neural networks has become increasingly pressing in the d

After rigorous testing, our joint AI project with Daiwa Securities is entering the full-scale production phase. We're bringing our agentic A…

AgentsDGX agent

After rigorous testing, our joint AI project with Daiwa Securities is entering the full-scale production phase. We're bringing our agentic AI systems to @Daiwa_JP’s wealth management teams to accelera

ETA: A New Agentic Paradigm for Embodied Tasks

AgentsDGX agent

arXiv:2608.03924v1 Announce Type: new Abstract: When will robots have their ChatGPT moment? Such a breakthrough requires a general-purpose robot that can handle unfamiliar tasks in unfamiliar environm

FLARE: Few-shot Learning-based Adaptive Reflective Engine

Model ReleasesDGX agent

arXiv:2608.02919v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in complex, compound AI systems where performance hinges on the quality of prompts. Recent state-

Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release (Ivan Mehta/TechCrunch)

Model ReleasesDGX agent

Ivan Mehta / TechCrunch: Hark, founded by Figure AI CEO Brett Adcock, previews Handoff, a computer use agent it says outperforms GPT-5.4 and Opus 4.8, and plans for a summer release — Hark, a startup

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-to…

SafetyDGX agent

Just had to create an 'accidental-cyberattacks' tag on my blog We're up to four now: the original OpenAI+Hugging Face one, Anthropic's me-too attacks, then two new ones from the UK AI Safety Institute

LFM2.5-2.6B on a OnePlus 13 at 17 tok/s ~ Pure CPU

Model ReleasesDGX agent

As you all know the model is 2.69B parameters with a 128K context window and purpose-built for multi-step agent workflows. What you are seeing is the Q4_K_M GGUF running on my own inference engine bui

Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation

Model ReleasesDGX agent

arXiv:2608.03577v1 Announce Type: new Abstract: Automation of Translation Quality Estimation (QE) has emerged as a widely discussed approach to managing translation quality at scale, and a growing num

Mistral introduces Shieldstral to provide lightweight policy-aware moderation for AI models

Model ReleasesDGX agent

French artificial intelligence startup Mistral AI SAS today introduced a lightweight multimodal safety artificial intelligence open-weight model that can classify outputs for AI models that outperform

Oilbird: Training-Free Speculative Decoding with Keys the Verifier Already Computes

Model ReleasesDGX agent

arXiv:2608.03839v1 Announce Type: new Abstract: Training-free speculative decoding drafts by matching an exact suffix of the context against a pool of earlier context. That lookup misses correct draft

Policy Fragmentation or Institutional Alignment? Institutional Governance of AI in Universities and Business Schools

SafetyDGX agent

arXiv:2608.03584v1 Announce Type: new Abstract: Artificial intelligence (AI) is rapidly transforming high-skilled domains, requiring higher education institutions (HEI) to balance the teaching of foun

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI…

SafetyDGX agent

Some people are surprised that APIs (aka what Anthropic, OpenAI, and others provide) are treated differently than open weights in the new AI model framework. I'm not surprised at all, and it's actuall

Sources: Google is in talks with AI coding agent startup Mechanize on a possible deal, potentially worth $1.5B+, to hire some of its talent and license its tech (Business Insider)

AgentsDGX agent

Business Insider: Sources: Google is in talks with AI coding agent startup Mechanize on a possible deal, potentially worth $1.5B+, to hire some of its talent and license its tech — Google wants its AI

4 Aug 2026

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Model ReleasesDGX agent

First of all, obviously I took some help from AI to type this post and this is the topic that enabled me to accomplish all that: https://old.reddit.com/r/LocalLLaMA/comments/1veow4b/deepseek_v4flash_2

← Previous
1…8788899091…169
Next →