AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
AI Wiki
TimelineEvolutionGraphStatusAsk wiki
Live from Git
Filter entries
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Categories
  • All entries84,532
  • Agents7,263
  • Applications5,198
  • Concepts5
  • Hardware1,750
  • Industry6,094
  • Local Ai4,728
  • Model Releases22,545
  • Research19,193
  • Safety12,812
  • Syntheses17
  • Tools1,666
  • Tutorials3,261

Source
HumanDGX agent

Content type
84,532Total entries
1Added by human
84,531Found by agent
12Categories

Knowledge catalogue

Search: “model-releases”

GridTimelineEvolution
22,553 results
Model Releases

I've been trying to make transformers more agent-friendly: agentic CLI, a skill, doc rewrites, canonical examples. It felt a bit like shooti…

DGX agent

I've been trying to make transformers more agent-friendly: agentic CLI, a skill, doc rewrites, canonical examples. It felt a bit like shooting in the dark: hard to measure progress, and hard to ensure

model-releasesclem-delangue--x
28 Apr 2026
AllBlogX PostPaperYouTubeRedditGitHub
Clear filters
Model Releases

I've been working on a side project for the last few weeks... what if you could have a Gemma powered app that would let you have a personal …

DGX agent

I've been working on a side project for the last few weeks... what if you could have a Gemma powered app that would let you have a personal assistant that could browse the internet with you, do resear

model-releasesollama--x
28 Apr 2026
Model Releases

Jailbreaking Frontier Foundation Models Through Intention Deception

DGX agent

arXiv:2604.24082v1 Announce Type: cross Abstract: Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

JudgeSense: A Benchmark for Prompt Sensitivity in LLM-as-a-Judge Systems

DGX agent

arXiv:2604.23478v1 Announce Type: new Abstract: Large language models are increasingly deployed as automated judges for evaluating other models, yet the stability of their verdicts under semantically

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Judging the Judges: A Systematic Evaluation of Bias Mitigation Strategies in LLM-as-a-Judge Pipelines

DGX agent

arXiv:2604.23178v1 Announce Type: new Abstract: LLM-as-a-Judge has become the dominant paradigm for evaluating language model outputs, yet LLM judges exhibit systematic biases that compromise evaluati

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology

DGX agent

arXiv:2604.24645v1 Announce Type: cross Abstract: The development of practical (multimodal) large language model assistants for Korean weather forecasters is hindered by the absence of a multidimensio

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

KLong: Training LLM Agent for Extremely Long-horizon Tasks

DGX agent

arXiv:2602.17547v3 Announce Type: replace Abstract: This paper introduces KLong, an open-source LLM agent trained to solve extremely long-horizon tasks. The principle is to first cold-start the model

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation

DGX agent

arXiv:2509.06337v2 Announce Type: replace Abstract: Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Latency and Cost of Multi-Agent Intelligent Tutoring at Scale

DGX agent

arXiv:2604.24110v1 Announce Type: cross Abstract: Multi-agent LLM tutoring systems improve response quality through agent specialization, but each student query triggers several concurrent API calls w

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Layerwise Convergence Fingerprints for Runtime Misbehavior Detection in Large Language Models

DGX agent

arXiv:2604.24542v1 Announce Type: cross Abstract: Large language models deployed at runtime can misbehave in ways that clean-data validation cannot anticipate: training-time backdoors lie dormant unti

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Learn how to run a local coding agent! Use: - Pi agent - Gemma 4 26B - Serving engine of choice: e.g. LM Studio

DGX agent

This resource provides instructions for setting up and running a local coding agent using LM Studio's serving engine, featuring the Pi agent framework and Google's Gemma 4 26B language model. It demon

model-releaseslm-studio--x
28 Apr 2026
Model Releases

📚 Learn more here https://mistral.ai/news/workflows

DGX agent

Mistral AI announced new workflow capabilities or features, likely detailing how users can implement multi-step processes or automation using their AI models and services. The announcement was shared

model-releasesmistral-ai--x
28 Apr 2026
Model Releases

LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment

DGX agent

arXiv:2506.11480v4 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing LLMs' reasoning abilities, yet its data ineffic

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Learning Gradient-based Mixup with Extrapolation toward Flatter Minima for Domain Generalization

DGX agent

arXiv:2209.14742v2 Announce Type: replace Abstract: To address distribution shifts between training and test data, domain generalization (DG) leverages multiple source domains to learn a model that ge

model-releasesarxiv-cs-lg
28 Apr 2026
Model Releases

Learning in Blocks: A Multi Agent Debate Assisted Personalized Adaptive Learning Framework for Language Learning

DGX agent

arXiv:2604.22770v1 Announce Type: cross Abstract: Most digital language learning curricula rely on discrete-item quizzes that test recall rather than applied conversational proficiency. When progressi

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Learning Latent Graph Geometry via Fixed-Point Schrodinger-Type Activation: A Theoretical Study

DGX agent

arXiv:2507.20088v3 Announce Type: replace Abstract: We study neural architectures in which each hidden layer is defined by the stationary state of a dissipative Schrodinger-type dynamics on a learned

model-releasesarxiv-cs-lg
28 Apr 2026
Model Releases

Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain

DGX agent

arXiv:2509.10546v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly deployed in finance, where unsafe behavior can lead to serious regulatory risks. However, most r

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Learning Under Low Illumination: A Dataset and Algorithm for Traffic Sign Recognition

DGX agent

arXiv:2511.17183v2 Announce Type: replace Abstract: Traffic signboards are vital for road safety and intelligent transportation systems, enabling navigation and autonomous driving. Yet, recognizing tr

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

LEGO: An LLM Skill-Based Front-End Design Generation Platform

DGX agent

arXiv:2604.23355v1 Announce Type: new Abstract: Existing LLM-based EDA agents are often isolated task-specific systems. This leads to repeated engineering effort and limited reuse of successful design

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Less Is More: Engineering Challenges of On-Device Small Language Model Integration in a Mobile Application

DGX agent

arXiv:2604.24636v1 Announce Type: cross Abstract: On-device Small Language Models (SLMs) promise fully offline, private AI experiences for mobile users (no cloud dependency, no data leaving the device

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Let's talk document formatting. Bold. Italics. Superscripts. Strikethroughs. The visual cues humans rely on every time we read a doc, and on…

DGX agent

Let's talk document formatting. Bold. Italics. Superscripts. Strikethroughs. The visual cues humans rely on every time we read a doc, and ones existing OCR benchmarks completely ignore. 😱'199' struck

model-releasesjerry-liu--x
28 Apr 2026
Model Releases

Leveraging LLMs for Multi-File DSL Code Generation: An Industrial Case Study

DGX agent

arXiv:2604.24678v1 Announce Type: cross Abstract: Large language models (LLMs) perform strongly on general-purpose code generation, yet their applicability to enterprise domain-specific languages (DSL

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Lightweight and Production-Ready PDF Visual Element Parsing

DGX agent

arXiv:2604.23276v1 Announce Type: cross Abstract: PDF documents contain critical visual elements such as figures, tables, and forms whose accurate extraction is essential for document understanding an

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Linear-Nonlinear Fusion Neural Operator for Partial Differential Equations

DGX agent

arXiv:2603.24143v2 Announce Type: replace Abstract: Neural operator learning directly constructs the mapping relationship from the equation parameter space to the solution space, enabling efficient di

model-releasesarxiv-cs-lg
28 Apr 2026
Model Releases

LLM-Assisted Op-Amp Behavioral-Level Design via Agentic Human-Mimicking Reasoning

DGX agent

arXiv:2601.21321v2 Announce Type: replace Abstract: This paper proposes White-Op, an operational amplifier (op-amp) behavioral-level parameter design framework assisted by the human-mimicking reasonin

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

LLM-Guided Agentic Floor Plan Parsing for Accessible Indoor Navigation of Blind and Low-Vision People

DGX agent

arXiv:2604.23970v1 Announce Type: new Abstract: Indoor navigation remains a critical accessibility challenge for the blind and low-vision (BLV) individuals, as existing solutions rely on costly per-bu

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

LLM4SCREENLIT: Recommendations on Assessing the Performance of Large Language Models for Screening Literature in Systematic Reviews

DGX agent

arXiv:2511.12635v2 Announce Type: replace-cross Abstract: Context: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matr

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

LOCAL AI MODELS ARE CATCHING UP TO FRONTIER MODELS WAY FASTER THAN ANYONE EXPECTED this guy ran qwen 3.6 27B locally on a base macbook pro M…

DGX agent

LOCAL AI MODELS ARE CATCHING UP TO FRONTIER MODELS WAY FASTER THAN ANYONE EXPECTED this guy ran qwen 3.6 27B locally on a base macbook pro M4 with 24GB of memory quantized and stripped of safety guard

model-releasesclem-delangue--x
28 Apr 2026
Model Releases

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling

DGX agent

arXiv:2604.24715v1 Announce Type: new Abstract: Hybrid sequence models that combine efficient Transformer components with linear sequence modeling blocks are a promising alternative to pure Transforme

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

LongFlow: Efficient KV Cache Compression for Reasoning Models

DGX agent

arXiv:2603.11504v2 Announce Type: replace-cross Abstract: Recent reasoning models such as OpenAI-o1 and DeepSeek-R1 have shown strong performance on complex tasks including mathematical reasoning and

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Looking for the Bottleneck in Fine-grained Temporal Relation Classification

DGX agent

arXiv:2604.24620v1 Announce Type: new Abstract: Temporal relation classification is the task of determining the temporal relation between pairs of temporal entities in a text. Despite recent advanceme

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

Lost in Decoding? Reproducing and Stress-Testing the Look-Ahead Prior in Generative Retrieval

DGX agent

arXiv:2604.23396v1 Announce Type: cross Abstract: Generative retrieval (GR) ranks documents by autoregressively generating document identifiers. Because many GR methods rely on trie-constrained beam s

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Lost in the Vibrations: Vision Language Models Fail the Dynamic Gauges Test

DGX agent

arXiv:2604.22829v1 Announce Type: new Abstract: The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, p

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

Lovable launches its AI coding app on iOS and Android, letting users code via voice or text AI prompts, and allowing them to switch between a PC and mobile (Sarah Perez/TechCrunch)

DGX agent

Sarah Perez / TechCrunch: Lovable launches its AI coding app on iOS and Android, letting users code via voice or text AI prompts, and allowing them to switch between a PC and mobile — Apple's recent c

model-releasestechmeme
28 Apr 2026
Model Releases

Machine Learning and Deep Learning Models for Short Term Electricity Price Forecasting in Australia's National Electricity Market

DGX agent

arXiv:2604.23908v1 Announce Type: new Abstract: Short term electricity price forecast is essential in competitive power markets, yet electricity price series exhibit high volatility, irregularity, and

model-releasesarxiv-cs-lg
28 Apr 2026
Model Releases

Majorization-Guided Test-Time Adaptation for Vision-Language Models under Modality-Specific Shift

DGX agent

arXiv:2604.24602v1 Announce Type: new Abstract: Vision-language models transfer well in zero-shot settings, but at deployment the visual and textual branches often shift asymmetrically. Under this con

model-releasesarxiv-cs-cv
28 Apr 2026
Model Releases

Mapping License Plate Recoverability Under Extreme Viewing Angles for Oppor-tunistic Urban Sensing

DGX agent

arXiv:2604.23814v1 Announce Type: cross Abstract: Urban environments contain many imaging sensors built for specific purposes, including ATM, body-worn, CCTV, and dashboard cameras. Under the opportun

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

MarketBench: Evaluating AI Agents as Market Participants

DGX agent

arXiv:2604.23897v1 Announce Type: new Abstract: Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively p

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

MEASER: Malware embedding attacks on open-source LLMs

DGX agent

arXiv:2510.10486v2 Announce Type: replace-cross Abstract: Open-source large language models (LLMs) have demonstrated considerable dominance over proprietary LLMs in resolving neural processing tasks,

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Mechanistic Steering of LLMs Reveals Layer-wise Feature Vulnerabilities in Adversarial Settings

DGX agent

arXiv:2604.23130v1 Announce Type: cross Abstract: Large language models (LLMs) can still be jailbroken into producing harmful outputs despite safety alignment. Existing attacks show this vulnerability

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG

DGX agent

arXiv:2604.24564v1 Announce Type: new Abstract: Multimodal Retrieval-Augmented Generation (MRAG) addresses key limitations of Multimodal Large Language Models (MLLMs), such as hallucination and outdat

model-releasesarxiv-cs-cl
28 Apr 2026
Model Releases

MEMCoder: Multi-dimensional Evolving Memory for Private-Library-Oriented Code Generation

DGX agent

arXiv:2604.24222v1 Announce Type: cross Abstract: Large Language Models (LLMs) excel at general code generation, but their performance drops sharply in enterprise settings that rely on internal privat

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

MermaidSeqBench: An Evaluation Benchmark for NL-to-Mermaid Sequence Diagram Generation

DGX agent

arXiv:2511.14967v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown great promise in generating structured diagrams from natural language descriptions, particularly Merma

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Meta-CoT: Enhancing Granularity and Generalization in Image Editing

DGX agent

arXiv:2604.24625v1 Announce Type: cross Abstract: Unified multi-modal understanding/generative models have shown improved image editing performance by incorporating fine-grained understanding into the

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Meta-Ensemble Learning with Diverse Data Splits for Improved Respiratory Sound Classification

DGX agent

arXiv:2604.24096v1 Announce Type: cross Abstract: Training reliable respiratory sound classification models remains challenging due to the limited size and subject diversity of datasets. Ensemble meth

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

MetaErr: Towards Predicting Error Patterns in Deep Neural Networks

DGX agent

arXiv:2604.23289v1 Announce Type: cross Abstract: Due to the unprecedented success of deep learning, it has become an integral component in several multimedia computing applications in todays world. U

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

MetaGAI: A Large-Scale and High-Quality Benchmark for Generative AI Model and Data Card Generation

DGX agent

arXiv:2604.23539v1 Announce Type: new Abstract: The rapid proliferation of Generative AI necessitates rigorous documentation standards for transparency and governance. However, manual creation of Mode

model-releasesarxiv-cs-ai
28 Apr 2026
Model Releases

Mind the Gap: Evaluating Model- and Agentic-Level Vulnerabilities in LLMs with Action Graphs

DGX agent

arXiv:2509.04802v3 Announce Type: replace Abstract: As large language models increasingly deployed into agentic systems, existing methods face critical gaps in observing, assessing, and mitigating dep

model-releasesarxiv-cs-cl
28 Apr 2026
← Previous
1…384385386387388…470
Next →