NeurionForgeAI Studio
STATION 01 — AI STUDIO
100% Local Inference & Fine-Tuning

Your Own AI Studio &Fine-Tuner,Built From Scratch.

Run open-weight LLMs locally via quantized GGUF execution, train custom LoRA/QLoRA adapters on your data, and orchestrate local coding agents — zero cloud dependencies.

The Toolkit

Everything AI Studio ships with today — and what's being built next.

LiveBuildingPlanned
Live

Local Inference Engine

Token-by-token WebSocket streaming with GGUF quantization, prompt caching, and sub-200ms time-to-first-token.

Live

PEFT / LoRA Studio

Dataset ingestion, background fine-tuning pipelines, real-time loss tracking, and GGUF adapter merging.

Live

Global Storage Pool

Unified D:/models repository shared across all local tooling, so nothing gets duplicated on disk.

Building

OpenAI-Compatible API

A drop-in local endpoint for any OpenAI-SDK app — same request shape, no API key, no code changes.

Building

Model Library & Quantizer

Pull, convert, and requantize open-weight checkpoints to GGUF directly from the UI — no terminal needed.

Building

Chat Playground

Multi-session chat with live adapter switching, system-prompt presets, and token-level telemetry.

Planned

Agent Mode

An opencode-style, tool-calling agent with real terminal access — running entirely against your local models.

Planned

Resource Monitor

Live VRAM, RAM, and GPU utilization per loaded model, so you know what fits before you load it.

The Build Path

Three phases, shipping in order — nothing is deferred past its phase.

Phase 0Complete

Foundations

  • Local inference engine baseline
  • CLI LoRA fine-tuning verification
  • Global D:/models storage pool
Phase 1Complete

Inference Studio

  • Streaming chat & session manager
  • HuggingFace Hub model downloader
  • Live RAM & CPU diagnostic HUD
Phase 2Active

Fine-Tuning Studio

  • JSONL dataset manager & validator
  • Real-time WebSocket loss telemetry
  • LoRA adapter library & studio testing
>_Inference baseline: Qwen2.5-1.5B-Instruct (Q4_K_M)
Local Core Offline
FastAPI Core + Next.js App Router