Technology

Models & post-training

Open models are the foundation. Ivertiq’s post-training is aimed at system capability — shaping the language model so the agentic layer can complete vertical workflows under customer SOPs, on sovereign appliances — not at improving the model as a standalone chatbot.

Why post-training is part of the product

Prompting and RAG alone rarely produce production-grade agents for regulated or ERP-heavy work. Those systems must call tools in the right order, verify external claims before accepting “success,” and stay efficient enough to run on appliance GPUs with real concurrency.

Ivertiq therefore treats post-training as a first-class layer of the Full-Stack AI Harness — alongside orchestration, validation, and vertical apps — not as a one-off science project.

Post-training objective: system capability

Post-training objectives are set holistically. The goal is not a better chatbot or a higher general benchmark; it is a model that helps the agentic layer complete system objectives under customer SOPs — on sovereign hardware, with provenance and human gates.

That means we train for the capabilities the workflow needs: domain judgment and the thinking and tool-use patterns that let agents plan, call skills, and verify outcomes. Domain knowledge without tool discipline fails in production; tool calling without vertical judgment fails the same way.

How those objectives are weighted is vertical- and workflow-specific. The mix is derived from what the Full-Stack system must reliably do — not from maximizing language-model capabilities in isolation.

Model progress is a tailwind

Frontier and near-frontier model advances — including rapid open releases — are a tailwind, not a threat. Ivertiq does not bet the product on one proprietary API: we assemble, customize, and optimize systems holistically for regulated verticals. New breakthroughs expand the toolkit. Better foundational models give richer dimensions for targeted, data-curated post-training toward vertical specialization.

Measured gains remain workload- and Validation Pack–specific; we do not claim that any single public release automatically delivers a fixed quantitative lift for every customer. The durable work is digesting new models and the methods they introduce, then deploying them under sovereignty and human-in-the-loop controls.

Example — ERP agentic workflows

Capabilities the system needs from the model

When ERP workflows require chain-of-thought and tool use — for example writing Python, retrieving information, or calling predictors for production yields — post-training is designed so the language model can assist the agentic layer, not only hold more ERP facts.

1. Domain knowledge

ERP / SOP language and procedure-aware judgment the agent needs to choose the right path: orchestration patterns, bilingual operational phrasing, and behavior generic base models do not own.

2. Thinking (CoT) efficiency

Reliable multi-step planning for those workflows — dense enough to be correct, efficient enough for appliance concurrency — not open-ended “think longer” as a general chat upgrade.

3. Tool use & verification

Call and verify the tools the system owns (code, connectors, predictors) with structured plans and anti–false-verify discipline — so agentic processes actually complete under human gates.

These three are an ERP-shaped illustration of system-derived training targets. Other verticals (life sciences, forensic review) set a different mix from their workflows and Validation Packs.

How it fits the stack

From post-trained model to governed workflow

Ivertiq Core

Local LLMs/SLMs, knowledge engine, and inference optimization — including quantized serving tuned for appliance form factors.

Harness Control

Governed agent runtime with skills/tools, policies, provenance, and human-in-the-loop gates that constrain what the model may do.

Vertical Apps

ERP intelligence, life-science assist, forensic review — workflows that consume the post-trained orchestrator rather than a generic chat endpoint.

Post-training shapes the model for system objectives. The Harness governs execution. Apps define the workflows those objectives serve.

Illustrative direction (ERP orchestration)

One vertical where system objectives drive the training mix: internal post-training programs on open dense models (for example Qwen-class 27B orchestrators) target ERP agent behavior — decompose multi-step work, call tools, and verify outcomes so the agentic layer can finish the job. Design goals include:

  • Much shorter first-turn thinking on ERP validation sets versus the unmodified base — reducing latency before the first tool call (illustrative internal results have shown large reductions in thinking overhead).
  • Faster path to tool use so agents spend GPU time acting, not narrating.
  • Stronger tool-calling and verification discipline for read/report/approval-style operations in bilingual (EN/ZH) settings.
  • Explicit non-goal: not a general MMLU/chat upgrade — a domain orchestration specialist that serves on-prem agentic workflows.

Detailed metrics, datasets, and eval protocols are shared with design partners and investors under NDA — not as public leaderboard theater.

Buyer outcomes

What this means in practice

Appliance fit

Efficient reasoning + right-sized models make Single-GB10 / Dual-GB10 deployments practical for departmental concurrency — without metering every token to a public API.

Agent reliability

System-oriented post-training is what turns “chat about ERP” into “plan the steps, call the skill, verify the result, wait for human confirm.”

Compounding moat

Each vertical’s traces, Validation Packs, and failure modes feed the next training loop — assets wrappers cannot copy from a prompt template.

What we do not claim

  • That post-training replaces human approval on regulated or financial actions
  • That every base-model capability improves equally (domain specialists can trade general benchmarks)
  • That post-training is primarily about leaderboard or chatbot scorecards — our objective is system capability under SOPs
  • That public cloud frontier models are “bad” — they are simply the wrong default for many sovereign workflows