InstructLab — Redesigning LLM Fine-Tuning for Open-Source Contributors

InstructLab is an open-source project from IBM and Red Hat that lets contributors improve large language models through structured knowledge and skill submissions. As an independent concept exploration, the studio redesigned its contributor experience — taming a deep taxonomy tree with Miller Columns, adding a Beginner/Expert mode toggle, and prototyping directly to code with the Carbon Design System and Vercel v0.

Role: UX & Front-End Design

Overview

InstructLab lets non-experts improve large language models by contributing structured knowledge and skills — no ML engineering required, at least in theory. The studio saw an opportunity to close the gap between that promise and the day-to-day experience of actually using the platform, and took on a self-directed redesign of its contributor-facing interface: how people navigate InstructLab's taxonomy, submit examples, and understand what the model has and hasn't learned.

Challenge

InstructLab organizes every contribution — a fact, a skill, a Q&A example — inside a deep hierarchical taxonomy tree. That structure is what makes the platform scalable, but it's also where usability broke down: contributors had to understand and navigate a multi-level category system just to find where their knowledge belonged, before they could even start contributing. For a platform built to invite broad participation, taxonomy navigation was the first wall most people hit.

Research: Designing for Non-Technical Domain Experts

The redesign was grounded in three personas representing InstructLab's real target audience: domain experts, not developers. Jordan Ellis, a customer success manager who wants to teach an AI the product knowledge in her head. David Kumar, a bank VP of finance who won't write code but needs airtight compliance and audit trails. And Patricia Williams, a senior HR manager training an AI assistant to help 35 managers navigate performance reviews after a company merger — the persona used to stress-test the interface end to end.

Patricia's full journey was mapped — from uploading anonymized review examples, through defining cultural principles and legal guardrails, to piloting with three managers and rolling out to all 35 — across six phases. That mapping surfaced a consistent thread: her trust in the system rose and fell with how much the AI showed its work. Black-box processing states, vague error messages, and silence during model training all read as risk to someone accountable for legally sensitive HR guidance.

Key Design Decisions

  • Miller Columns for taxonomy navigation: a three-column picker that shows category, subcategory, and specific topic side by side, so contributors can see the full path as they drill down instead of losing context in nested menus.
  • Beginner and Expert modes: a step-by-step wizard with progress indicators for first-time contributors, and a single-page, all-fields view — built around Patricia's need to move fast without sacrificing the audit trail — for repeat contributors.
  • IBM Carbon Design System as the visual and component foundation, chosen for consistency with the platform's existing ecosystem and to keep the concept credible as a plausible evolution of the real product.
  • A YAML editor surfaced for power users (ML engineers, developers) who prefer direct, code-level control over a contribution's structure, without abandoning the guided flow for everyone else.
  • Vercel v0 for direct-to-code prototyping — generating production-quality React and Tailwind components from prompts, so the concept could be evaluated as working software rather than static comps.

From Persona to Interface

Patricia's journey map turned directly into interface decisions. Her anxiety during model training became step-labeled progress states ("Analyzing examples… Extracting patterns… Configuring guardrails…") instead of a bare spinner. Her fear of missing a legal risk became a guardrail-review screen that proactively surfaces protected categories and lets her add her own rules — like blocking the AI from ever suggesting termination language, a decision that belongs to HR alone. Her worry about opaque AI judgment became visible confidence signals on generated responses, so she could see when the model was uncertain rather than confidently wrong.

Design Targets (Simulated Walkthroughs)

Because this was a concept exploration rather than a live deployment, real user outcomes were not measured. What was defined, as part of the persona journey mapping, were the targets the design was built to hit — the bar used to judge whether a given screen or flow was working. For Patricia's scenario: greater than 75% manager adoption within 60 days of rollout, 100% redirect accuracy on legally protected topics, and at least two hours saved per manager per month. These are illustrative design goals from a simulated walkthrough, not measured results, and should be read that way.

Conclusion

InstructLab's core idea — that anyone with domain expertise should be able to shape a model's behavior — depends entirely on whether non-technical contributors can actually use the tool. This concept treats that as the design problem worth solving: a taxonomy that doesn't require a map, a workflow that flexes between guided and expert use, and an interface that shows its reasoning to people whose job is to catch the moment it's wrong.