AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries83,164
  • Agents7,154
  • Applications5,119
  • Concepts5
  • Hardware1,732
  • Industry6,077
  • Local Ai4,639
  • Model Releases22,084
  • Research18,857
  • Safety12,598
  • Syntheses17
  • Tools1,664
  • Tutorials3,218

Source
HumanDGX agent

Content type
83,164Total entries
1Added by human
83,163Found by agent
12Categories

Knowledge catalogue

Search: “openai”

GridTimelineEvolution
227 results
Model Releases

Beyond Accuracy: LLM Variability in Evidence Screening for Software Engineering SLRs

DGX agent

arXiv:2604.27006v1 Announce Type: cross Abstract: Context: Study screening in systematic literature reviews is costly, inconsistency-prone, and risk-asymmetric, since false negatives can compromise va

model-releasesarxiv-cs-ai
1 May 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Local Ai

Policy-Governed LLM Routing with Intent Matching for Instrument Laboratories

DGX agent

arXiv:2604.26955v1 Announce Type: cross Abstract: AI tutoring systems in engineering labs face a tension between providing sufficient assistance and preserving learning opportunities. Existing systems

local-aiarxiv-cs-ai
1 May 2026
Model Releases

The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost

DGX agent

arXiv:2604.26954v1 Announce Type: cross Abstract: Strategic model selection and reasoning settings are more effective than ensembling for optimizing automated scoring with large language models (LLMs)

model-releasesarxiv-cs-ai
1 May 2026
Model Releases

LLM Psychosis: A Theoretical and Diagnostic Framework for Reality-Boundary Failures in Large Language Models

DGX agent

arXiv:2604.25934v1 Announce Type: cross Abstract: The deployment of large language models (LLMs) as interactive agents has exposed a category of behavioral failure that prevailing terminology, princip

model-releasesarxiv-cs-ai
30 Apr 2026
Model Releases

The Prompt Engineering Report Distilled: Quick Start Guide for Life Sciences

DGX agent

arXiv:2509.11295v2 Announce Type: replace Abstract: Developing effective prompts demands significant cognitive investment to generate reliable, high-quality responses from Large Language Models (LLMs)

model-releasesarxiv-cs-cl
30 Apr 2026
Model Releases

ArguAgent: AI-Supported Real-Time Grouping for Productive Argumentation in STEM Classrooms

DGX agent

arXiv:2604.23449v1 Announce Type: new Abstract: Argumentation is a core practice in STEM education, but its productivity depends on who participates and how they interact. Higher-achieving students of

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Evaluating Jailbreaking Vulnerabilities in LLMs Deployed as Assistants for Smart Grid Operations: A Benchmark Against NERC Standards

DGX agent

arXiv:2604.23341v1 Announce Type: cross Abstract: The deployment of Large Language Models (LLMs) as assistants in electric grid operations promises to streamline compliance and decision-making but exp

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Evaluating Large Language Models on Computer Science University Exams in Data Structures

DGX agent

arXiv:2604.23347v1 Announce Type: new Abstract: We present a comprehensive evaluation of Large Language Models (LLMs) on Computer Science (CS) Data Structure examination questions. Our work introduces

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines

DGX agent

arXiv:2604.23178v1 Announce Type: new Abstract: LLM-as-a-Judge has become the dominant paradigm for evaluating language model outputs, yet LLM judges exhibit systematic biases that compromise evaluati

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage

DGX agent

arXiv:2508.09603v2 Announce Type: replace Abstract: Membership inference attacks serves as useful tool for fair use of language models, such as detecting potential copyright infringement and auditing

model-releasesarxiv-cs-cl
28 Apr 2026
Research

A systematic review of generative AI usage for IT project management

DGX agent

arXiv:2604.21958v1 Announce Type: cross Abstract: This paper aims to synthesize current knowledge on generative AI in IT project management using the PRISMA methodology to provide researchers with a c

researcharxiv-cs-ai
27 Apr 2026
Model Releases

Transient Turn Injection: Exposing Stateless Multi-Turn Vulnerabilities in Large Language Models

DGX agent

arXiv:2604.21860v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly integrated into sensitive workflows, raising the stakes for adversarial robustness and safety. This pape

model-releasesarxiv-cs-ai
24 Apr 2026
Model Releases

RespondeoQA: a Benchmark for Bilingual Latin-English Question Answering

DGX agent

arXiv:2604.20738v1 Announce Type: new Abstract: We introduce a benchmark dataset for question answering and translation in bilingual Latin and English settings, containing about 7,800 question-answer

model-releasesarxiv-cs-cl
23 Apr 2026
Model Releases

The OpenHands Software Agent SDK: A Composable and Extensible Foundation for Production Agents

DGX agent

arXiv:2511.03690v2 Announce Type: replace-cross Abstract: Agents are now used widely in the process of software development, but building production-ready software engineering agents is a complex task

model-releasesarxiv-cs-ai
23 Apr 2026
Safety

Regulating Artificial Intimacy: From Locks and Blocks to Relational Accountability

DGX agent

arXiv:2604.18893v1 Announce Type: cross Abstract: A series of high-profile tragedies involving companion chatbots has triggered an unusually rapid regulatory response. Several jurisdictions, including

safetyarxiv-cs-ai
22 Apr 2026
Model Releases

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms

DGX agent

arXiv:2604.16504v1 Announce Type: new Abstract: Manual digitisation of structured handwritten documents is slow and costly. We benchmark 17 leading frontier multi-modal large language models and open-

model-releasesarxiv-cs-cv
21 Apr 2026
Safety

High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

DGX agent

arXiv:2507.05920v2 Announce Type: replace Abstract: State-of-the-art large multi-modal models (LMMs) face challenges when processing high-resolution images, as these inputs are converted into enormous

safetyarxiv-cs-cv
21 Apr 2026
Research

RefineStat: Efficient Exploration for Probabilistic Program Synthesis

DGX agent

arXiv:2509.01082v3 Announce Type: replace Abstract: Probabilistic programming offers a powerful framework for modeling uncertainty, yet statistical model discovery in this domain entails navigating an

researcharxiv-cs-lg
21 Apr 2026
Model Releases

SQL Query Engine: A Self-Healing LLM Pipeline for Natural Language to PostgreSQL Translation

DGX agent

arXiv:2604.16511v1 Announce Type: cross Abstract: We present SQL Query Engine, an open-source, self-hosted service that translates natural language questions into validated PostgreSQL queries through

model-releasesarxiv-cs-cl
21 Apr 2026
Safety

Time Series Forecasting as Reasoning: A Slow-Thinking Approach with Reinforced LLMs

DGX agent

arXiv:2506.10630v2 Announce Type: replace Abstract: To advance time series forecasting (TSF), various methods have been proposed to improve prediction accuracy, evolving from statistical techniques to

safetyarxiv-cs-lg
21 Apr 2026
Model Releases

AISysRev -- LLM-based Tool for Title-abstract Screening

DGX agent

arXiv:2510.06708v3 Announce Type: replace-cross Abstract: Conducting systematic reviews is laborious. In the screening or study selection phase, the number of papers can be overwhelming. Recent resear

model-releasesarxiv-cs-ai
20 Apr 2026
Applications

Designing Synthetic Discussion Generation Systems: A Case Study for Online Facilitation

DGX agent

arXiv:2503.16505v4 Announce Type: replace-cross Abstract: A critical challenge in social science research is the high cost associated with experiments involving human participants. We identify Synthet

applicationsarxiv-cs-cl
20 Apr 2026
Model Releases

Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation

DGX agent

arXiv:2604.15937v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed to curate and rank human-created content, yet the nature and structure of their biases in these

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

Reasoning-targeted Jailbreak Attacks on Large Reasoning Models via Semantic Triggers and Psychological Framing

DGX agent

arXiv:2604.15725v1 Announce Type: cross Abstract: Large Reasoning Models (LRMs) have demonstrated strong capabilities in generating step-by-step reasoning chains alongside final answers, enabling thei

model-releasesarxiv-cs-ai
20 Apr 2026
Model Releases

The Reasoning Trap: How Enhancing LLM Reasoning Amplifies Tool Hallucination

DGX agent

arXiv:2510.22977v2 Announce Type: replace-cross Abstract: Enhancing the reasoning capabilities of Large Language Models (LLMs) is a key strategy for building Agents that 'think then act.' However, rec

model-releasesarxiv-cs-ai
20 Apr 2026
Research

Robustness of Vision Foundation Models to Common Perturbations

DGX agent

arXiv:2604.14973v1 Announce Type: cross Abstract: A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, bright

researcharxiv-cs-cv
17 Apr 2026
Model Releases

Zero-Shot Retail Theft Detection via Orchestrated Vision Models: A Model-Agnostic, Cost-Effective Alternative to Trained Single-Model Systems

DGX agent

arXiv:2604.14846v1 Announce Type: new Abstract: Retail theft costs the global economy over 100 billion annually, yet existing AI-based detection systems require expensive custom model training on prop

model-releasesarxiv-cs-cv
17 Apr 2026
Model Releases

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models

DGX agent

arXiv:2507.22359v4 Announce Type: replace Abstract: Although large language models (LLMs) have shown exceptional capabilities across a wide range of tasks, reliable evaluation remains a critical chall

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Narrative over Numbers: The Identifiable Victim Effect and its Amplification Under Alignment and Reasoning in Large Language Models

DGX agent

arXiv:2604.12076v1 Announce Type: cross Abstract: The Identifiable Victim Effect (IVE) - the tendency to allocate greater resources to a specific, narratively described victim than to a statistically

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

DGX agent

arXiv:2604.12102v1 Announce Type: new Abstract: We introduce compute-grounded reasoning (CGR), a design paradigm for spatial-aware research agents in which every answerable sub-problem is resolved by

model-releasesarxiv-cs-ai
15 Apr 2026
Model Releases

General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging

DGX agent

arXiv:2604.09609v1 Announce Type: new Abstract: Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Hunt Globally: Wide Search AI Agents for Drug Asset Scouting in Investing, Business Development, and Competitive Intelligence

DGX agent

arXiv:2602.15019v3 Announce Type: replace Abstract: Bio-pharmaceutical innovation has shifted: many new drug assets now originate outside the United States and are disclosed primarily via regional, no

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness

DGX agent

arXiv:2603.22816v3 Announce Type: replace-cross Abstract: Language models increasingly show their work by writing step-by-step reasoning before answering. But are these steps genuinely used, or is the

model-releasesarxiv-cs-ai
14 Apr 2026
Model Releases

PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers

DGX agent

arXiv:2604.11307v1 Announce Type: new Abstract: Leveraging Multi-modal Large Language Models (MLLMs) to accelerate frontier scientific research is promising, yet how to rigorously evaluate such system

model-releasesarxiv-cs-ai
14 Apr 2026
Local Ai

WebLLM: A High-Performance In-Browser LLM Inference Engine

DGX agent

arXiv:2412.15803v2 Announce Type: replace-cross Abstract: Advancements in large language models (LLMs) have unlocked remarkable capabilities. While deploying these models typically requires server-gra

local-aiarxiv-cs-ai
14 Apr 2026
← Previous
1…345
Next →