How to Use Jev to Build a Market Making Engine

Within 48 hours of Jev shipping, a trading loop was already reading an order book and making typed quoting decisions at sub-second latency, fast enough to post limit orders continuously within the execution cycle. The community did not theorize about what this enables. It just built.
A market making engine is a decision machine. Jev is a decision model. Every decision a market maker makes continuously, whether to quote, how wide to set the spread, whether incoming flow is informed, is a typed question with a calibrated probability attached. Jev was built to answer exactly that shape of question, at exactly the speed a trading system can use, with exactly the output type an execution layer can consume without any translation step in between.
This article builds the full architecture. Three decision layers, complete schema design for each, a fallback architecture that makes it safe to run, and the full engine wired end to end.
I am Ruuj, a backend developer working on system design, researcher, and quantitative trading systems. For collaborations and discussions, DMs are open.
Let's get started
Chapter 1: Why Jev and Market Making Were Built for Each Other
Most AI development has converged on smarter chat and deeper reasoning. TypeSafe went in a completely different direction
Jev is designed as a System One Model, a category built for fast structured decisions that software can act on directly. The name is a reference to Daniel Kahneman's framework: System 1 is the fast instinctive response, System 2 is the slow deliberate one. Large language models are System 2, producing outputs through sequential autoregressive decoding rather than a single parallel pass. Jev is System 1, a single parallel pass that returns a typed answer with a calibrated confidence score attached.
The mechanics are precise. You pass it unstructured program state alongside a typed schema defining the possible answers. Jev evaluates everything in one parallel pass and returns a typed object whose fields are the decision variables your code reads directly, with calibrated probabilities on every answer. There is no string to parse, no value to extract, no validation step before the answer is usable. The output is natively consumable by software in a single line.
Three question primitives cover every decision shape a market making engine needs.
Choice selects one option from a defined set and returns a calibrated probability on each option. Score returns a calibrated value on a defined numeric scale. Noul, Jev's binary primitive, returns a yes or no with a calibrated probability on each side rather than a single binary value.
Every decision built across this article is one of these three shapes.
The property that makes all of this matter for trading is calibration. A calibrated confidence score reflects the empirical probability of the decision being correct, a property that should be validated against live production data before being relied upon in sizing decisions. A 0.9 confidence score on a QUOTE decision is not just a high number. It is a quantitative input the next layer in the system can act on proportionally. That distinction is the foundation everything else in this article rests on.
Why market making fits this architecture exactly
A market making engine makes three classes of decision continuously across every cycle. Whether to be quoting or stepping back from the market. How wide the spread should be and which direction to skew the quotes to manage inventory. Whether incoming flow is informed or uninformed.
Every one of these is a typed question with a discrete answer and a calibrated probability attached. All of them require a decision fast enough to live inside the execution cycle. All of them map exactly onto the primitives Jev returns.
The Avellaneda Stoikov framework makes the architecture explicit. The optimal reservation price for a market maker adjusts away from the midprice by:
r = m − q × γ × σ² × (T − t)
Where r is the reservation price, m is the current midprice, q is current inventory, γ is the market maker's risk aversion parameter, σ² is return variance, and T − t is time remaining in the session. The reservation price skews away from midprice in the direction that reduces inventory, and by an amount that grows with both the size of the inventory and the time pressure of the session approaching its end.
The optimal spread around that reservation price is:
δ = γ × σ² × (T − t) + (2/γ) × ln(1 + γ/κ)
Where κ is the decay parameter governing how rapidly market order arrival intensity falls as quotes are placed further from the midprice. The spread's first term widens with volatility, inventory risk, and time remaining. The second term reflects order arrival structure and is independent of time. Every variable in both expressions is observable state. The challenge is that a parametric formula weights those variables according to fixed parameters calibrated at a point in time. Jev weights them according to calibrated intelligence derived from the full context of the current market state, including context the formula cannot encode.
The three decision layers built across the next three Parts map directly onto these expressions and extend them with something no parametric formula alone provides: a calibrated judgment about what the current market conditions actually require.
Once you see market making as a sequence of typed probabilistic decisions derived from observable state, and once you understand that Jev is a model whose output natively matches that exact shape, the architecture that follows becomes the obvious way to build it.
The quoting layer determines whether to be in the market at all. That is where the build starts.
Chapter 2: The Quoting Layer
The first decision every market making engine makes on each cycle is whether to be in the market at all. This seems like the simplest decision in the engine. It is actually the most consequential.
A market maker quoting when it should be stepped back pays the full adverse selection cost of every informed trade that hits it while it is standing there. That cost is paid in full at the moment of execution. The subsequent spread earned on uninformed flow does not straightforwardly recover it.
Most market making systems handle this with a rules table. Volatility above a threshold: step back. Spread above a threshold: step back. The rules table is precise and auditable, and it has one specific limitation worth understanding clearly. The conditions that should trigger a step back are context dependent in ways a fixed threshold cannot capture. A volatility spike during a macro release is a structurally different signal from the same volatility spike during thin overnight trading. A sharp order book imbalance during a known auction window reads differently from the same imbalance during normal continuous trading. The rules table treats both identically because it operates on individual variables in isolation. Jev weighs all available context simultaneously and returns a calibrated probability on the decision rather than a binary output from a single threshold applied to a single variable.
Schema design for the quoting layer
Schema design is the craft that separates a well-functioning intelligence layer from an expensive noise generator. Every field passed to Jev in the quoting layer is expressed as a relative measure, a multiple or fraction of its own rolling baseline, so inputs remain comparable across different market conditions and regimes. Absolute values tie inputs to specific levels of variables and degrade calibration when conditions shift. Relative values let the same schema function coherently across a low volatility trending session and a high volatility event-driven one without any recalibration.
Order book state grounds the schema. Current spread expressed as a multiple of its rolling baseline tells Jev whether conditions are tight or wide relative to recent history rather than in absolute terms. Depth at each level passed as a fraction of rolling average depth at that level communicates whether the book is full or thin relative to normal. The ratio of bid depth to offer depth across the top several levels captures directional imbalance in the passive liquidity structure, which often develops before it appears in executed trades.
Recent trade flow gives Jev the active side of the market. Directional imbalance over a short lookback window expressed as the fraction of volume that was buyer initiated tells Jev whether flow has been consistently one directional or balanced. Average trade size as a fraction of average daily volume at this time of session distinguishes normal sized flow from oversized orders. Trade frequency in the most recent short window relative to the rolling baseline captures whether market activity is accelerating.
Realized volatility passes short window volatility as a multiple of the longer window baseline. A multiple well above one means current conditions are materially more volatile than recent history. Framing this as a multiple rather than an absolute level allows Jev to understand the degree of deviation from normal without being tied to any specific volatility level.
Inventory state passes current inventory as a fraction of the session position limit and its directional sign. The quoting decision changes meaningfully when inventory approaches its limit in either direction, and Jev needs that context to make a calibrated judgment about whether quoting at current parameters is appropriate.
Session context passes time elapsed as a fraction of total session length. Microstructure shifts systematically across the trading day and a quoting decision at the open carries different weight from the same decision during peak-liquidity hours even when the raw order book state looks similar.
Jev receives all five categories in a single state object and returns a Choice between QUOTE and STEP BACK with a calibrated probability on each.
How calibrated probability feeds forward
The probability on QUOTE carries more information than a binary yes or no, and the architecture is designed to use it fully.
High confidence QUOTE at stable conditions means quote at the baseline spread. Lower confidence QUOTE means quote and widen proportionally, using the degree of uncertainty as a direct scaling input to spread width. The quoting decision's confidence travels forward into the spread layer rather than being discarded after producing a direction. Uncertainty in the quoting layer automatically produces a more conservative spread posture in the next layer without requiring a separate rule to enforce it.
result = jev.decide(state=quoting_state, questions=quoting_schema)
if result.decision == "STEP_BACK" and result.confidence > threshold:
execution.pull_quotes()
else:
spread_layer.run(base_multiplier = 1 + (1 - result.confidence))The quoting layer tells the engine whether to be in the market and exactly how confident that decision is. That confidence does not stop there. It travels forward as a quantitative input into the spread layer, which receives it, uses it, and combines it with inventory state to determine how wide to quote and in which direction to skew.
That is the next decision.
Chapter 3 The Spread and Inventory Layer
Once the quoting layer returns QUOTE, two decisions happen simultaneously. How wide the spread should be. Which direction to skew the quotes to manage inventory risk.
These are related but structurally distinct problems and the engine handles them as separate typed questions evaluated in the same Jev call.
Spread width is a function of volatility, the adverse selection component of recent flow, and the confidence carried forward from the quoting layer. Inventory skew is a function of the current position, the urgency of reducing it, and the cost of holding it through the remainder of the session.
Getting either wrong is expensive in ways that compound differently. A spread too tight in high adverse selection conditions gifts P&L to informed traders one execution at a time and the cost is paid immediately. A spread too wide in benign conditions gives up volume to competitors quoting tighter and the cost accumulates across the session. Inventory not managed progressively through the day becomes inventory that must be unwound near the close at worse prices, under pressure the engine created by its own earlier inaction.
Spread width as a Score primitive
The spread layer schema passes the quoting layer confidence score as its first input, short window volatility as a multiple of the longer baseline, the current spread across available venues as a multiple of its rolling baseline, and the adverse selection component of recent trades measured by post trade midprice drift on the same side. Jev returns a Score on a defined scale representing the spread multiple relative to the configured baseline. The execution layer multiplies that score directly against the baseline spread and sets the quotes. One typed number consumed directly as a price with no intermediate step.
Inventory skew as Choice and Score
The inventory schema passes current inventory as a fraction of the session limit, the trajectory of the position over recent cycles showing whether it is growing or reducing, time remaining in the session as a fraction of total session length, and the estimated cost of carrying the current position to end of session.
Jev returns a Choice between SKEW BID, SKEW OFFER, and NEUTRAL, plus a Score representing the magnitude of the skew needed. The reservation price moves in the direction that encourages the market to take the inventory, by exactly the amount the Score indicates.
This is the Avellaneda Stoikov reservation price adjustment operationalized as a typed Jev output. The expression q × γ × σ² × (T − t) provides the theoretical structure: skew the reservation price in proportion to inventory size, risk aversion, volatility, and time remaining. Jev provides the calibrated version of that adjustment, weighing all four variables simultaneously in the context of current market conditions including conditions the parametric formula cannot encode: the trajectory of recent position changes, the cross-venue spread environment, and the confidence state carried forward from the quoting decision.
result = jev.decide(state=spread_inventory_state, questions=spread_inventory_schema)
quotes = execution.build_quotes(
spread = baseline_spread * result.spread_multiple,
skew = result.skew_magnitude * direction(result.skew_direction)
)Spread and inventory are the engine's decisions about its own risk, derived from internal state it can observe and control. The hardest decision comes from outside the engine entirely. Whether the flow arriving right now is informed or uninformed determines whether every decision made so far in this cycle was worth making.
That is the adverse selection layer.
Chapter 4 | The Adverse Selection Layer
Adverse selection is the defining risk of market making.
The spread earned from uninformed flow is precisely what covers the losses from informed flow. The market making engine that cannot distinguish between the two is structurally exposed to a risk it cannot price and cannot see until the P&L reflects it.
The widely cited academic framework for estimating informed trading probability is VPIN, the Volume Synchronized Probability of Informed Trading, developed by Easley, Lopez de Prado, and O'Hara. It estimates the fraction of volume originating from informed traders by classifying trades as buyer initiated or seller initiated and measuring directional imbalance across volume synchronized buckets rather than time synchronized ones. Volume synchronization matters because informed traders tend to act when volume is present, so time-bucket measures can miss clustering that volume-bucket measures capture.
VPIN as a rolling aggregate has a specific production limitation worth understanding. It reflects a lagged statistical window across recent trades rather than the specific characteristics of the individual order arriving right now. It produces a single number that requires a separate mapping to a trading decision rather than a typed decision with a calibrated probability directly attached.
Jev operationalizes the same underlying measurement as a typed probabilistic decision on the current order's specific context. The conceptual foundation is the same. The output is native, typed, and calibrated.
What the adverse selection layer schema carries
Five fields describe the informational content of the incoming order and its immediate context.
Order size as a multiple of rolling average order size at this time of session captures the degree to which this order is disproportionately large relative to normal flow. Informed traders tend to size positions more deliberately than uninformed ones, and large relative size is one of the most consistent signals in the microstructure literature.
Order aggression describes whether the order is immediately marketable or was placed passively and became marketable as the price moved to it. A marketable order crossing the spread reflects a different intent from a passive order that sat and waited, and that distinction carries meaningful signal about the order's purpose in the market.
The pattern of cancellations and resubmissions in the order book in the window immediately preceding this order captures whether the passive liquidity structure was cleared before the incoming order arrived. This sequence is more consistent with strategic informed positioning than with normal uninformed market participation.
Midprice drift following recent executions on the same side as this order measures whether similar recent orders produced systematic subsequent price movement. Post trade drift on the same side is the clearest observable signature of informed flow in the microstructure literature and the variable VPIN is fundamentally designed to capture in aggregate form.
Cross-venue spread conditions at the exact moment of this order's arrival capture whether the order is routed to the venue with the tightest spread, a routing pattern more consistent with informed flow optimizing for best execution than with uninformed flow seeking liquidity wherever it is available.
Jev returns a Choice between INFORMED and UNINFORMED with a calibrated probability on each, plus a Score representing the urgency of the response given all the evidence in the schema simultaneously.
Proportional response driven by calibrated confidence
The response is proportional to the calibrated confidence rather than binary, and that proportionality is what separates this layer from a threshold-based rules check.
High confidence INFORMED at high urgency Score triggers either an immediate spread widen on the hit side scaled to the Score, or a full step back through the quoting layer path if the Score exceeds the configured threshold. Intermediate confidence triggers an asymmetric adjustment: the side being hit widens by an amount proportional to the Score while the other side remains competitive to preserve volume from uninformed flow arriving alongside it. Low confidence maintains current quotes and routes the order to the fallback audit log.
result = jev.decide(state=flow_state, questions=adverse_selection_schema)
if result.flow_type == "INFORMED":
if result.response_urgency > step_back_threshold:
execution.pull_quotes()
else:
execution.widen_hit_side(multiplier = 1 + result.response_urgency)A fixed threshold produces a binary output: above it widen, below it hold. Calibrated confidence produces a continuous response that costs less in lost volume than a binary widen and protects better than a binary hold across the full range of conditions the engine encounters in live operation.
Three layers. A quoting decision that determines whether to be in the market. A spread and inventory layer that determines how wide and which direction. An adverse selection layer that classifies incoming flow and responds proportionally to calibrated confidence.
Wire all three together and the engine thinks continuously. That is the final Part.
Chapter 5: The Full Engine
The full engine runs the three layers with one architectural property that changes its performance profile entirely.
The spread layer and the adverse selection layer run simultaneously rather than sequentially because Jev evaluates all typed questions in a single parallel pass. The quoting layer fires first. If it returns STEP BACK the cycle ends. If it returns QUOTE the spread schema and the adverse selection schema both go to Jev in the same call, evaluated in parallel, returned together. The outputs combine to produce the final quote in one pass.
Three schema design principles
Schema design is the most consequential engineering decision in the entire system. Three principles govern every choice made across all three layers.
Relative values over absolute values. Every field is expressed as a multiple or fraction of its own rolling baseline. This makes inputs comparable across regimes and allows calibration to hold across different market conditions without requiring the schema to be rebuilt every time the market enters a different environment. A volatility multiple of two means the same thing whether absolute volatility is low or high: current conditions are twice as volatile as recent history. That stability is what makes the intelligence generalize.
Context over raw signals. Midprice drift following recent executions on the same side carries more informational content about adverse selection than raw trade imbalance. Spread as a multiple of its own baseline carries more about current market stress than the absolute spread level. Contextualising each input against its own recent history is what produces calibrated decisions rather than reactive ones that respond to the absolute level of a variable rather than its deviation from normal.
Forward-flowing confidence. The probability output of the quoting layer becomes a scaling input to the spread layer. The adverse selection Score determines the magnitude of the execution response. Confidence that flows forward makes the three layers a connected decision architecture rather than three independent checks that happen to run in sequence. Each layer knows not just what it decided but how certain it was, and that certainty shapes what the next layer does.
The fallback architecture
Every layer has a defined fallback when Jev returns confidence below a configured threshold. The quoting layer falls back to a conservative baseline spread and quotes. The spread layer falls back to baseline spread without skew. The adverse selection layer maintains current quotes and routes the order to the audit log. Each layer degrades independently so a low confidence result on one layer does not affect the others.
Every fallback is logged with the full state object that produced it. The fallback log is the primary input for schema improvement over time. Patterns in the log reveal which state representations are producing uncertain decisions and which fields need to be reframed or enriched. A practitioner reviews the log on a defined cadence, modifies underperforming field definitions, and redeploys. Kill switches and hard position limits enforced at the execution layer are the minimum additional safeguards required before running this in production.
What changes architecturally
Every previous approach to adding AI intelligence to a market making system placed that intelligence alongside the execution cycle. A text generating model producing a recommendation, a parsing layer extracting a value from the output, a rule acting on the extracted value. The intelligence influenced decisions but was always separated from the execution layer by intermediate steps.
Jev places the intelligence inside the cycle. The typed output of each layer flows directly into the next layer's input or directly into execution logic. The output of a Jev call is a typed object whose fields the execution layer reads in a single line.
The gap between adjacent and inside is the entire architectural argument for building this.
The engine is complete. Three layers wired together, each speaking the same language, each producing typed decisions the next layer consumes directly, each improving its calibration from the production evidence the fallback log accumulates over time.
The Summary
Every decision in a market making engine has always been a typed question with a calibrated probability attached. Whether to quote. How wide. Whether the flow arriving right now changes everything. The missing piece was never the question. It was a model whose output matched the shape of the answer natively, at the speed and in the format the execution layer could consume without translation.
That gap closed this week.
The intelligence in this architecture does not sit beside the execution cycle advising it. It sits inside the cycle as the cycle. The quoting decision is a typed probabilistic answer acted on directly. The spread is a calibrated Score multiplied into a price. The adverse selection response is a Choice with a confidence score executed proportionally in the same pass. There is no intermediate step between the intelligence and the action it produces.
The schema built here gets sharper from its own operation. The fallback log generated from the first week of running is a precise record of where the intelligence needs to improve, and that record compounds into a materially different engine over time.







