Local review draft Not a launched service

Independent AI research & development

Less to generate.
More to deliver.

Strong models do the hard reasoning. Small models turn compact answers into complete results.

We’re exploring learned compression between the two, so delivering an answer takes less time without sacrificing what matters.

Explore the research

Working prototype. Learned codec under development.

Same intent. A smaller message.
From rich reasoning to a compact message and a full answer Multiple lines converge into a short central representation, then expand into complete output. Conceptual illustration, not measured compression.
ReasonEncodeExpand
Conceptual view. Actual compression and fidelity are measured together.

The approach

Keep the reasoning.
Shorten the handoff.

Generating a full answer is expensive. We’re investigating whether a strong model can communicate its solution through a shorter representation that a fast model learns to reconstruct.

  1. 01

    Learn to compress

    Train an encoder and decoder on validated answers. Reward smaller messages only when reconstruction preserves the required meaning and behavior.

  2. 02

    Connect to strong reasoning

    Train a prompt controller to guide a frozen frontier model toward representations the decoder understands. Test that connection throughout training.

  3. 03

    Measure the whole answer

    Compare correctness and total completion time, including the controller, strong model and decoder. Fewer tokens alone aren’t a successful result.

First evidence · September 2026

A promising result.
A specific experiment.

On one regex-engine coding task, our prompted Astra → Spark pipeline passed the same checks as Astra alone, with a lower median completion time.

This prototype uses implementation plans and exact core code. It does not yet use a trained codec. Two runs of one task do not establish general reliability or speed.

Read experiment details

Median generation time Lower is better

Astra high128.19 s
Astra → Spark91.97 s
Full-task correctness, two runs per path
PathPassed
Astra high2 / 2
Spark low0 / 2
Astra → Spark2 / 2

82 check groups per candidate. Timing includes generation pipeline overhead; independent test execution is excluded.

The company

AugmentOps.
Built around a research question.

We’re a bootstrapped software startup developing AI-output compression technology. Our current work focuses on the training and evaluation of learned encoders, fast decoders and prompt controllers.

The prototype is experimental. We’re working toward a system that delivers complete, correct answers faster, and testing where that approach holds up.

augmentops@goldeneyetools.fr

goldeneyetools.fr AugmentOps’ web address