ROT

OTIF Calculator

OTIF is the single most widely used supplier scorecard metric in retail because it's binary and unforgiving by design: an order that arrives on time but short a case doesn't count, and an order that's complete but three days late doesn't count either. You end up with the OTIF percent and the distance from a 95 percent target. The distance is the part that starts the supplier conversation.

Reviewed by Bhanu PrakashLast updated August 10, 2026
Inputs

Enter your numbers

orders
orders
Result

Your calculation

OTIF

86.00%

Gap to 95% Target

9.00 pts

On-Time In-Full Orders

430

Total Orders

500

Formula Used

(Orders On Time and In Full ÷ Total Orders) × 100

Formula

How the number is calculated

(Orders On Time and In Full ÷ Total Orders) × 100

The "AND" in OTIF is the entire point of the metric. An order can score well on either dimension individually and still fail OTIF entirely: 98 percent on-time delivery combined with 90 percent fill-rate completeness does not average to some blended 94 percent OTIF. If those two failure modes hit different orders, actual combined OTIF could be as low as 88 percent (only orders that were both on-time and complete count), and if they overlap on the same orders it could be higher. The two conditions must be measured jointly, on the same order, not calculated as two separate percentages and averaged. OTIF can be measured at the order level (did this entire PO arrive on-time and complete) or the line level (did this individual SKU line within the PO arrive on-time and complete). Line-level OTIF is stricter and typically a better operational indicator, since a large multi-line PO can be "order-level OTIF" if the bulk of lines arrive fine even while a specific SKU consistently short-ships, a pattern that order-level measurement can mask for months. OTIF is a lagging indicator of supplier reliability, not a forecasting tool. A supplier's trailing 12-month OTIF trend is far more useful for vendor negotiation and sourcing decisions than any single-period snapshot, since single periods can be distorted by one-off disruptions (weather, port congestion, a single bad batch) that don't reflect the supplier's underlying operational discipline.

Worked Example

Of 500 supplier orders received in a quarter, 430 were delivered on time and complete. OTIF = (430 ÷ 500) × 100 = 86 percent, below the typical 95 percent benchmark and worth escalating. Now the what-ifs. Breaking down the 70 failed orders: 45 were late but complete, and 25 were on-time but short-shipped. Fixing only the lateness issue (say, through better carrier selection) would improve OTIF to (430+45)/500 = 95 percent, hitting benchmark, while fixing only the completeness issue would improve it to (430+25)/500 = 91 percent, still below benchmark. This decomposition is critical for supplier conversations: "your OTIF is 86 percent" is far less actionable than "your OTIF is 86 percent, driven mostly by late delivery, not fill-rate gaps," which points the supplier at the specific process to fix. Next quarter, the same supplier improves lateness significantly (only 10 late orders now) but completeness issues persist (still 25 short-shipped): OTIF = (500-10-25)/500 = 93 percent, a meaningful improvement but still short of the 95 percent target, and the remaining gap is now concentrated entirely in fill-rate performance, which typically points to a warehouse-side or inventory-planning problem on the supplier's end rather than a logistics/carrier problem. Finally, compare line-level versus order-level measurement on the same data: if those 25 short-shipped orders were each missing just one line item out of an average 8 lines per order, line-level OTIF (evaluating each of the roughly 4,000 total lines individually) would show a different, typically lower, percentage than order-level OTIF, because line-level measurement penalizes every partial short-ship rather than only counting whole-order failures.

Frequently Asked Questions

What is a healthy OTIF target?+

95 percent or higher is the standard target for top-tier suppliers in most retail categories. Below 90 percent typically signals a supplier relationship that needs active management. Below 85 percent, the cost of stockouts, expediting, and emergency replenishment becomes significant enough to warrant formal supplier escalation or dual-sourcing consideration.

Is OTIF measured at line or order level?+

Both views have value, but they answer different questions. Order-level OTIF is simpler to track and communicate. Line-level OTIF is stricter and usually a better leading indicator of underlying supplier process problems, since it catches partial short-ships that order-level measurement can mask. Track both if resources allow; if choosing one, line-level provides more diagnostic value.

How should OTIF failures be decomposed for supplier conversations?+

Split every failure into "late but complete" versus "on-time but incomplete" versus "both late and incomplete." Each failure type points to a different root cause on the supplier's side (carrier/logistics issues versus warehouse/inventory issues versus systemic capacity problems), and a supplier scorecard conversation grounded in this breakdown is far more actionable than a single aggregate percentage.

How does OTIF relate to reorder point and safety stock?+

Directly. A supplier's OTIF performance feeds the lead-time variability term in the safety stock formula; unreliable suppliers (low or volatile OTIF) require more safety stock to protect the same service level, which ties up more working capital. Improving supplier OTIF is one of the highest-leverage ways to reduce safety stock without accepting more stockout risk.

What time window should OTIF be measured over?+

Trailing 90 days is a common standard for ongoing supplier scorecards, long enough to smooth single-incident noise but recent enough to reflect current performance. Trailing 12 months is more appropriate for annual sourcing and contract renewal decisions, since it captures seasonal patterns a 90-day window might miss.

Should OTIF weight orders by dollar value or count them equally?+

Depends on the purpose. Simple order-count OTIF (as this calculator computes) is easiest to track and communicate broadly. Dollar-weighted OTIF gives more analytical weight to high-value orders, which is often more relevant for assessing actual business impact, since a failure on a $50,000 order matters more than a failure on a $500 order even though both count equally in a simple percentage.

What's considered "in full" when partial substitutions are allowed?+

This needs to be defined explicitly in the vendor agreement before measurement begins. Some retailers count an approved substitution as "in full" if quantity and category match; others count only exact-SKU fulfillment as in-full. Ambiguity here is a common source of disputes during supplier scorecard reviews, so the definition should be written into the contract, not inferred after the fact.

How quickly should a supplier below target be escalated?+

Two consecutive measurement periods below target (e.g., two consecutive months below 90 percent) is a common trigger for formal escalation. A single bad period is often noise; a sustained trend below target reflects a real operational problem worth addressing directly, potentially including dual-sourcing critical SKUs while the issue is resolved.

Related Articles

Deep-dive guides that explain the math behind this calculator.

Related Calculators

Explore Related Resources

Handpicked benchmarks, templates and guides to help you dig deeper.

Five core calculators every buyer, merchandiser and category manager reads together. Open the metric that is behind, and let the others sanity-check it.