AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “research”

GridTimelineEvolution
25,879 results
Agents

Mitigating Context Interference for Reliable and Efficient Search Agents

DGX agent

arXiv:2608.10743v1 Announce Type: new Abstract: Recent research empowers Large Language Models (LLMs) as multi-turn search agents to iteratively retrieve and generate outputs until complex tasks are s

agentsarxiv-cs-cl
12 Aug 2026
Model Releases
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters

New Muse-Glimmer-30B SoTA Quants - hopefully a new lineup :)

DGX agent

Hey Folks, I've been making quants for a while - recently I took a short break to get into hardcore research (submitted my first EMNLP paper during it!). Along the way, I built up a little arsenal of

model-releasesr-localllama
12 Aug 2026
Model Releases

On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation

DGX agent

arXiv:2608.11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years. However, existing research has largely focused on English-only settin

model-releasesarxiv-cs-ai
12 Aug 2026
Safety

ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions

DGX agent

arXiv:2608.10621v1 Announce Type: new Abstract: Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formula

safetyarxiv-cs-lg
12 Aug 2026
Agents

Recovering Wasted Compute in Autoresearch Agents

DGX agent

arXiv:2608.10424v1 Announce Type: new Abstract: A slew of recent works develop agents for solving research problems end-to-end, a paradigm increasingly referred to as autoresearch. Such agents have in

agentsarxiv-cs-ai
12 Aug 2026
Model Releases

Seq2Synth: Benchmarking Temporal Fidelity in Synthetic Sequential Tabular Data

DGX agent

arXiv:2607.15606v2 Announce Type: replace Abstract: Synthetic sequential tabular data are increasingly used for privacy-preserving data sharing and data-driven research, but evaluating their fidelity

model-releasesarxiv-cs-lg
12 Aug 2026
Safety

Toward a Theory of Value in AI Alignment

DGX agent

arXiv:2608.10327v1 Announce Type: new Abstract: Can AI systems be aligned to human values? The popularization of large language models (LLMs) and multi-modal foundation models has seen a rise in harms

safetyarxiv-cs-ai
12 Aug 2026
Applications

Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing

DGX agent

arXiv:2608.09289v1 Announce Type: new Abstract: Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates thes

applicationsarxiv-cs-cl
11 Aug 2026
Safety

Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy

DGX agent

arXiv:2608.09857v1 Announce Type: cross Abstract: Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on

safetyarxiv-cs-ai
11 Aug 2026
Model Releases

Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models

DGX agent

arXiv:2608.08822v1 Announce Type: new Abstract: Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and b

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

Autorubric: A Unifying Framework for Rubric-Based LLM Evaluation on Non-Verifiable Tasks

DGX agent

arXiv:2603.00077v3 Announce Type: replace-cross Abstract: Rubric-based LLM judges have become indispensable for evaluating and optimizing systems on non-verifiable tasks, where success cannot be reduc

model-releasesarxiv-cs-ai
11 Aug 2026
Applications

Bridging the Gap Between Semantics and Reconstruction:Unifying Sign Language Translation and Production

DGX agent

arXiv:2608.09045v1 Announce Type: cross Abstract: Recent advances in sign language (SL) research have shown a trend toward unifying multiple sign language understanding (SLU) subtasks, such as isolate

applicationsarxiv-cs-ai
11 Aug 2026
Model Releases

Efficient Human-Contact Representation for Human-Scene Interaction

DGX agent

arXiv:2608.09388v1 Announce Type: new Abstract: Human-scene interaction is an active research topic with several industrial applications in virtual reality, gaming, robotics, and surveillance. Despite

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our app…

DGX agent

Introducing 𝗘𝘅𝘁𝗿𝗮𝗰𝘁𝗕𝗲𝗻𝗰𝗵: the most comprehensive benchmark for information extraction from complex enterprise documents. Our applied research team tested: 14 systems — frontier VLMs, coding agents, ex

model-releasesjerry-liu--x
11 Aug 2026
Safety

Is CLIP Cross-Eyed? Revealing and Mitigating Center Bias in the CLIP Family

DGX agent

arXiv:2604.05971v2 Announce Type: replace-cross Abstract: Recent research has shown that contrastive vision-language models such as CLIP often lack fine-grained understanding of visual content. While

safetyarxiv-cs-cl
11 Aug 2026
Model Releases

Parameter Exploration for RLVR via Variational Learning

DGX agent

arXiv:2608.09805v1 Announce Type: cross Abstract: Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an importan

model-releasesarxiv-cs-ai
11 Aug 2026
Model Releases

PROSLEX: A Novel Dataset for Expert-Annotated Legal Statute Prediction for Indian Judiciary

DGX agent

arXiv:2608.08830v1 Announce Type: new Abstract: Legal Statute Prediction (LSP) involves automatically identifying relevant legal statutes given factual descriptions in legal documents, typically frame

model-releasesarxiv-cs-ai
11 Aug 2026
Tutorials

SPD Learn: A Geometric Deep Learning Python Library for Neural Decoding Through Trivialization

DGX agent

arXiv:2602.22895v2 Announce Type: replace-cross Abstract: Implementations of symmetric positive definite (SPD) matrix-based neural networks for neural decoding remain fragmented across research codeba

tutorialsarxiv-cs-lg
11 Aug 2026
Model Releases

The Cell Must Go On: Agar.io for Continual Reinforcement Learning

DGX agent

arXiv:2505.18347v3 Announce Type: replace-cross Abstract: Continual reinforcement learning (RL) concerns agents that are expected to learn continually, rather than converge to a policy that is then fi

model-releasesarxiv-cs-ai
11 Aug 2026
Applications

Towards Adaptive Super-Resolution and Quality Assessment via Test-Time Adaptation

DGX agent

arXiv:2608.08508v1 Announce Type: new Abstract: This paper presents doctoral research on adaptive video super-resolution and perceptual quality modeling under real-world conditions. Existing video sup

applicationsarxiv-cs-cv
11 Aug 2026
Model Releases

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery

DGX agent

arXiv:2608.08132v1 Announce Type: new Abstract: Reconstruction of 3D objects from a single image is a challenging research problem in computer vision. The key challenge is the lack of critical informa

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Zero-shot 2D Grounding with Novel Affordance Types

DGX agent

arXiv:2608.08929v1 Announce Type: new Abstract: 2D affordance grounding aims to locate the region of an object that a human can interact with. Existing research focuses on recognizing affordance types

model-releasesarxiv-cs-cv
11 Aug 2026
Model Releases

Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing

DGX agent

arXiv:2608.07437v1 Announce Type: new Abstract: Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate t

model-releasesarxiv-cs-ai
10 Aug 2026
Safety

How Should I Pick a Foundation Model for My Robot? In Favor of a Community Evaluation Framework for Social Robots

DGX agent

arXiv:2608.06898v1 Announce Type: cross Abstract: Researchers who seek to build social robot applications on foundation models are faced with a difficult question: how should we pick a model? Public l

safetyarxiv-cs-cl
10 Aug 2026
Safety

The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents

DGX agent

arXiv:2608.06663v1 Announce Type: new Abstract: Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour

safetyarxiv-cs-cl
10 Aug 2026
Model Releases

CyberKimi just dropped strong results on one of ExploitBench’s hardest V8 bugs , points away from Mythos

DGX agent

Hey everyone ! Quick share from the cyber + local LLM side of things that I found interesting. During this week’s hacker summer camp, an AI researcher and reverse malware engineer veteran 'lordx64' on

model-releasesr-ollama
9 Aug 2026
Model Releases

KLQ: Training-free measured rotation quantization. Beats all training-free rotation-based quantization methods on W4A4KV4-bits. Llama 3.2 1B KLQ-quantized beats SpinQuant and gets close to ReSpinQuant without GPTQ/LDLQ rounding.

DGX agent

First of all, I'm not a lab, this was a solo summer research project that finally culminated into the github repo and the writeup. The repo includes a much deeper dive with methods, findings about qua

model-releasesr-localllama
9 Aug 2026
Model Releases

SQLite compressed text-history prototypes

DGX agent

Research: SQLite compressed text-history prototypes I'm perennially interested in options for storing revision histories in relational databases. While out on a dog walk I had a new idea: how about ta

model-releasessimon-willison
9 Aug 2026
Local Ai

Quick survey (2 min) on trust in hardware specs for open-source models

DGX agent

Hi everyone, I'm a systems analysis student researching a problem a lot of you probably know well: how much you actually trust the published VRAM/RAM requirements for open-source models before trying

local-air-ollama
8 Aug 2026
Model Releases

We analyzed DeepSeek V4 Flash and GPT-5.6 Luna on DeepSWE. A DeepSeek-first cascade with test-suite verification solved MORE tasks than Luna…

DGX agent

Researchers from TogetherAI analyzed DeepSeek V4 Flash and GPT‑5.6 Luna on the DeepSWE benchmark. The study found that employing a DeepSeek‑first cascade with test‑suite verification solved more tasks

model-releasestogether-ai--x
8 Aug 2026
Agents

Agentic Software Issue Resolution with Large Language Models: A Survey

DGX agent

arXiv:2512.22256v2 Announce Type: replace-cross Abstract: Software issue resolution aims to address real-world issues in software repositories based on natural language descriptions provided by users,

agentsarxiv-cs-ai
7 Aug 2026
Model Releases

AI Playing Business Games: Benchmarking Large Language Models on Managerial Decision-Making in Dynamic Simulations

DGX agent

arXiv:2509.26331v2 Announce Type: replace Abstract: The rapid advancement of LLMs sparked significant interest in their potential to augment or automate managerial functions. One of the most recent tr

model-releasesarxiv-cs-ai
7 Aug 2026
Model Releases

Deep Generalised Mixed Models: a Novel Neural Network Structure for Analysing Hierarchical Data

DGX agent

arXiv:2608.05930v1 Announce Type: cross Abstract: The experience sampling method (ESM) is a longitudinal research design where participants report their thoughts, emotional states and behaviours multi

model-releasesarxiv-cs-lg
7 Aug 2026
Agents

How TReNDS automates root-cause analysis with Amazon Bedrock

DGX agent

TReNDS, a research center at Georgia State University, built an agentic AI pipeline on Amazon Bedrock and the open-source Strands Agents SDK that automatically investigates production errors in real t

agentsaws-ml-blog
7 Aug 2026
Tutorials

Measuring and Detecting Harmful AI Sycophancy

DGX agent

arXiv:2608.05624v1 Announce Type: new Abstract: Sycophantic responses are becoming pervasive in large language models (LLMs), and prior work has pointed out that some of them could be harmful. This pa

tutorialsarxiv-cs-ai
7 Aug 2026
Safety

Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges

DGX agent

arXiv:2405.15604v4 Announce Type: replace Abstract: Text generation has become more accessible than ever, and the growing interest in these systems, especially those using large language models, has s

safetyarxiv-cs-cl
7 Aug 2026
Local Ai

TLNM: Externally Validated Tooth Detection, Numbering and Segmentation from Smartphone Photographs Using Mask R-CNN

DGX agent

arXiv:2608.06275v1 Announce Type: new Abstract: Oral health issues affect billions globally, but the cost and limited access to professional dental care hinder preventive oral healthcare. Research rel

local-aiarxiv-cs-cv
7 Aug 2026
Agents

Tracing the Heart: An Evidence-Linked Pipeline for Heart-Failure Feature Engineering

DGX agent

arXiv:2608.06366v1 Announce Type: new Abstract: Electronic health record (EHR) feature engineering is a major bottleneck in clinical research and AI, accounting for 39-45% of data scientists' workload

agentsarxiv-cs-ai
7 Aug 2026
Model Releases

Unifying Structured and Unstructured Data Insights with BQ Search Innovations

DGX agent

Modern enterprises possess a vast amount of unstructured data, yet they frequently encounter significant challenges in managing and extracting value from it. Historically, unlocking the insights hidde

model-releasesgoogle-cloud-ai
7 Aug 2026
Agents

ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation

DGX agent

arXiv:2608.05026v1 Announce Type: cross Abstract: High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reli

agentsarxiv-cs-ai
6 Aug 2026
Agents

Eigenius: A Typed Knowledge-Graph DBMS with Epistemic Stratification and Institution-Mediated Reasoning

DGX agent

arXiv:2608.04457v1 Announce Type: cross Abstract: As 'AI Scientists' emerge to drive research via the Model Context Protocol (MCP), systems relying on ephemeral scripts will fail. The sheer scale of s

agentsarxiv-cs-ai
6 Aug 2026
Model Releases

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation

DGX agent

arXiv:2608.04374v1 Announce Type: cross Abstract: Large language models can produce fluent financial analysis, but fluency alone does not establish whether a report is suitable for institutional deliv

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

IslamicTurathBench: A Multi-Task, Multi-Discipline Benchmark for Evaluating Large Language Models on the Islamic Scholarly Tradition (turath)

DGX agent

arXiv:2608.04703v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where an

model-releasesarxiv-cs-cl
6 Aug 2026
Model Releases

K-EXAONE 2.0 Technical Report

DGX agent

arXiv:2608.04505v1 Announce Type: new Abstract: This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward glo

model-releasesarxiv-cs-cl
6 Aug 2026
Safety

NSF-HRPT: Neural Semantic Field meets Hierarchical Risk Perception Tree for Safety-Critical Scenario Assessment

DGX agent

arXiv:2608.04776v1 Announce Type: new Abstract: The ability to accurately assess and anticipate risks in safety-critical scenarios is crucial for autonomous driving systems. While existing research ha

safetyarxiv-cs-ai
6 Aug 2026
Local Ai

Promptable Animal Pose Tracking Across Species

DGX agent

arXiv:2608.04995v1 Announce Type: new Abstract: Animal pose estimation and tracking is important for wildlife monitoring and conservation research, and with limited expert time for labelling automated

local-aiarxiv-cs-cv
6 Aug 2026
Model Releases

SciCode-Verified: How Benchmark Defects Underestimated the Scientific-Coding Ability of Language Models

DGX agent

arXiv:2608.04975v1 Announce Type: cross Abstract: SciCode is the standard measure of the scientific-coding ability of language models: research-level problems that demand both frontier scientific theo

model-releasesarxiv-cs-ai
6 Aug 2026
Model Releases

Trident : How to Break Deep Reinforcement Learning Cyber Defenses (Agentic)

DGX agent

arXiv:2608.04317v1 Announce Type: cross Abstract: Autonomous cyber defense systems based on Deep Reinforcement Learning (DRL) have attracted significant research attention, yet remain evaluated almost

model-releasesarxiv-cs-ai
6 Aug 2026
← Previous
1…434435436437438…540
Next →