01Write the metric definition before calculating
Name the supplier, buying entity, site or lane, reporting period, products or categories, currency if value is used and measurement object. State whether the score is by PO line, schedule line, order, shipment, delivery or receipt. Define due population, event source, delivery window, accepted quantity, exclusions, responsibility rules, dispute deadline and version owner. Public and commercial contracts use different OTIF definitions; one UK procurement example measures customer order lines against a booking window and full delivery without specified claims. Treat it as proof that definitions must be explicit, not as a rule for every wholesale program. Keep target and formula separate so changing a service target does not silently change historical calculation logic.
02Freeze the denominator of due observations
Build the eligible population from approved purchase orders and delivery schedules due in the reporting period. Preserve PO, line, schedule, version, supplier, ship-from, destination, item, ordered quantity, unit and the governing date. Include overdue open lines according to the written rule. Exclude canceled or superseded demand only when approval predates the measurement cutoff and the history remains visible. Decide how blanket orders, call-offs, split schedules, consolidated deliveries and reopened lines behave. A supplier cannot be scored fairly from receipts alone because missing deliveries would disappear. Start from what was due, then trace each observation forward to despatch, transport and receipt. Reconcile denominator count to the open-PO and change-control records.
03Define the on-time test
Choose the promised event: ship date, carrier handover, requested arrival, booked delivery appointment, gate arrival or accepted receipt. Record time zone, calendar, cutoff and early/late tolerance. Use the contract or agreed order rule to decide whether requested, confirmed or approved revised date governs. Do not let a supplier confirmation unilaterally replace the buyer's required date. When the buyer approves a change, preserve old and new dates and approval time. Separate supplier production delay, buyer reschedule, carrier delay, customs event, dock congestion and receiving delay so responsibility can be analyzed without rewriting the event. A delivery can be on time for carrier arrival but late for accepted receipt; publish which one the metric uses.
04Define the in-full test
Use one controlled unit and conversion. Compare ordered or scheduled quantity with the quantity accepted under the written rule. State whether over-delivery passes, fails or is separately controlled; do not use excess units to offset a shortage on another line unless the commercial agreement allows it. Decide whether damaged, wrong-item, mislabeled, expired, quarantined or later-rejected units count as full. Keep received, inspected, accepted, rejected, returned and credited quantities distinct. For split deliveries, define whether the line passes only when cumulative accepted quantity reaches the threshold by the due window. Preserve lot, serial or package evidence where required. A full carton count is not full delivery when the item, unit or condition is wrong.
05Classify exclusions, data gaps and disputes
Create limited reason codes for approved buyer change, cancellation, documented force majeure where applicable, missing buyer access, duplicate line, data repair and other contract-defined exclusions. Each exclusion needs owner, evidence, effective date and scope. Report exclusions separately; never blend them into passes. Missing evidence should remain missing or pending, not automatically pass. Give the supplier a line-level statement showing due rule, events, quantity result and responsibility classification. Record dispute date, reason, evidence, reviewer, decision and metric version. Lock the original result and publish a controlled restatement when a dispute changes it. Monitor exclusion and dispute rates because a high OTIF score built on many exclusions is not a strong result.
06Calculate and reconcile the score
For a binary line-based model, set On time = 1 only when the chosen event falls within the defined window, In full = 1 only when accepted quantity meets the rule, and OTIF pass = 1 only when both are 1. Then calculate 100 × OTIF passes ÷ eligible due observations. Also publish on-time-only, in-full-only, fail-both and pending counts. Example: 82 eligible lines, 70 passing both, 5 late-only, 3 short-only and 4 failing both produce 85.37% OTIF; the bridge totals back to 82. Do not round intermediate results to force a target. Recalculate from source records and compare with the scorecard extract. If value-weighted or unit-weighted performance is needed, label it as a different metric rather than calling it the same OTIF rate.
07Use OTIF to improve sourcing decisions
Break results by supplier site, product family, lane, buyer, requested lead time, order-change frequency, failure reason and severity. Review both percentage and observation volume. A small supplier with two lines is not directly comparable to one with two thousand without context. Connect recurring late failures to acknowledgment, capacity, production status and recovery controls; connect short delivery to allocation, packaging, quantity tolerance and receiving accuracy. Set corrective actions with owner and due date, then measure whether the failure mode declines. Do not use OTIF as the only supplier decision: quality, compliance, responsiveness, cost, risk and data integrity remain separate dimensions. Retain the line-level cohort so a score can be reproduced after a later receipt or credit correction.