FOR TEAMS BUILDING WITH LANGUAGE MODELS

What is your AI
doing wrong?

Wrong answers. Unreliable agents. A model that works in the demo but struggles in the real workflow.

Tell us what your AI is doing wrong. Let’s have a conversation. We help teams find the failure, understand the model, and test what could make it better.

30-minute introduction · Choose a time on Calendly · No cost

Conceptual illustration of a layered neural model with diagnostic signals passing through it
OBSERVE → COMPARE → VERIFYProduct illustration
Make model behavior
an engineering question.
Reproducible evaluationPer-case evidenceInternal-state expertiseClear next steps

A FOCUSED ENGAGEMENT

One model problem.
A decision you can act on.

For teams dealing with unsupported answers, inconsistent task behavior, or a quality score they cannot confidently explain.

01 / ESTABLISH

Get a trustworthy baseline.

We review your rubric, labeled cases and saved outputs. Scoring mistakes and data overlap are checked before conclusions are drawn.

02 / DIAGNOSE

See the pattern in the failures.

Receive case-level findings, matched comparisons where available, and a ranked account of what deserves engineering attention.

03 / DECIDE

Choose the next controlled test.

Get a rerunnable evaluation package and a scoped plan for improvement. If model access supports it, that plan can include internal measurements and steering.

THE PROPRIOCEPTIVE APPROACH

Look inside.
Test the change.

Our research works with hidden-state measurements and causal interventions. In a separately scoped control pilot, we can instrument a supported model, apply a bounded change, and measure what happens downstream.

Every proposed improvement is compared with the unchanged model and relevant simpler controls. Helpful results, harmful results and null results all stay in the record.

Explore our public research ↗
YOUR WORKFLOWOne defined failure
DIAGNOSTICMeasure & explain
OPTIONAL CONTROL PILOTIntervene & compare
DELIVERYEvidence your team can rerun

START WITH THE PROBLEM

A conversation first.
A useful scope next.

Bring one example of what is going wrong. We’ll discuss the impact, your model access, and whether a focused engagement would help.

BEFORE YOU START

A few useful answers.

Can you work with an API-only model?

Yes, for the diagnostic using your evaluation code and saved inputs and outputs. Hidden-state inspection and steering require access to a supported model runtime and a separately agreed pilot.

What do we need to provide?

One workflow, a rubric, up to 100 labeled cases, saved model answers and any comparison outputs. We agree scope and a suitable transfer route before you share project materials. Use redacted or synthetic examples for the first conversation; we agree a secure transfer route before exchanging sensitive materials.

What could an engagement include?

A technical diagnostic and reusable evaluation artifacts. You receive the findings even if your baseline holds up or no useful intervention is identified. Implementation, certification, large-scale labeling and ongoing monitoring are separate.

What happens on the first call?

We discuss your workflow, a concrete failure, and what model or evaluation access you have. If there is a useful fit, we propose a defined scope, fee and next step. The introduction is free and does not commit you to a project.

Who delivers the work?

Proprioceptive AI, led by Logan Napolitano. Our public research documents the methods behind the work. Your engagement is private and findings are not published without your permission.

Bring us one failure
you need to understand.

Book a conversation →