Get a trustworthy baseline.
We review your rubric, labeled cases and saved outputs. Scoring mistakes and data overlap are checked before conclusions are drawn.
FOR TEAMS BUILDING WITH LANGUAGE MODELS
Wrong answers. Unreliable agents. A model that works in the demo but struggles in the real workflow.
Tell us what your AI is doing wrong. Let’s have a conversation. We help teams find the failure, understand the model, and test what could make it better.
30-minute introduction · Choose a time on Calendly · No cost

A FOCUSED ENGAGEMENT
For teams dealing with unsupported answers, inconsistent task behavior, or a quality score they cannot confidently explain.
We review your rubric, labeled cases and saved outputs. Scoring mistakes and data overlap are checked before conclusions are drawn.
Receive case-level findings, matched comparisons where available, and a ranked account of what deserves engineering attention.
Get a rerunnable evaluation package and a scoped plan for improvement. If model access supports it, that plan can include internal measurements and steering.
THE PROPRIOCEPTIVE APPROACH
Our research works with hidden-state measurements and causal interventions. In a separately scoped control pilot, we can instrument a supported model, apply a bounded change, and measure what happens downstream.
Every proposed improvement is compared with the unchanged model and relevant simpler controls. Helpful results, harmful results and null results all stay in the record.
Explore our public research ↗START WITH THE PROBLEM
Bring one example of what is going wrong. We’ll discuss the impact, your model access, and whether a focused engagement would help.
Choose a time to talk with Logan about your AI workflow.
Understand a recurring failure before committing to a larger engineering project.
Fixed scope and fee agreed after the introduction.
Explore the deliverable →For teams with access to a supported model runtime, explore internal measurements and controlled steering.
Model access, compute, success criteria and price are agreed together.
Discuss your model →BEFORE YOU START
Yes, for the diagnostic using your evaluation code and saved inputs and outputs. Hidden-state inspection and steering require access to a supported model runtime and a separately agreed pilot.
One workflow, a rubric, up to 100 labeled cases, saved model answers and any comparison outputs. We agree scope and a suitable transfer route before you share project materials. Use redacted or synthetic examples for the first conversation; we agree a secure transfer route before exchanging sensitive materials.
A technical diagnostic and reusable evaluation artifacts. You receive the findings even if your baseline holds up or no useful intervention is identified. Implementation, certification, large-scale labeling and ongoing monitoring are separate.
We discuss your workflow, a concrete failure, and what model or evaluation access you have. If there is a useful fit, we propose a defined scope, fee and next step. The introduction is free and does not commit you to a project.
Proprioceptive AI, led by Logan Napolitano. Our public research documents the methods behind the work. Your engagement is private and findings are not published without your permission.