IMS vs WMS Explained: The Sequencing Decision Most Retailers Get Wrong
The question is rarely which system to buy, it is which to buy first. A framework based on where the operational pain sits today.
OTIF is the single most widely used supplier scorecard metric in retail because it's binary and unforgiving by design: an order that arrives on time but short a case doesn't count, and an order that's complete but three days late doesn't count either. You end up with the OTIF percent and the distance from a 95 percent target. The distance is the part that starts the supplier conversation.
OTIF
86.00%
Gap to 95% Target
9.00 pts
On-Time In-Full Orders
430
Total Orders
500
Formula Used
(Orders On Time and In Full ÷ Total Orders) × 100
(Orders On Time and In Full ÷ Total Orders) × 100
The "AND" in OTIF is the entire point of the metric. An order can score well on either dimension individually and still fail OTIF entirely: 98 percent on-time delivery combined with 90 percent fill-rate completeness does not average to some blended 94 percent OTIF. If those two failure modes hit different orders, actual combined OTIF could be as low as 88 percent (only orders that were both on-time and complete count), and if they overlap on the same orders it could be higher. The two conditions must be measured jointly, on the same order, not calculated as two separate percentages and averaged. OTIF can be measured at the order level (did this entire PO arrive on-time and complete) or the line level (did this individual SKU line within the PO arrive on-time and complete). Line-level OTIF is stricter and typically a better operational indicator, since a large multi-line PO can be "order-level OTIF" if the bulk of lines arrive fine even while a specific SKU consistently short-ships, a pattern that order-level measurement can mask for months. OTIF is a lagging indicator of supplier reliability, not a forecasting tool. A supplier's trailing 12-month OTIF trend is far more useful for vendor negotiation and sourcing decisions than any single-period snapshot, since single periods can be distorted by one-off disruptions (weather, port congestion, a single bad batch) that don't reflect the supplier's underlying operational discipline.
Of 500 supplier orders received in a quarter, 430 were delivered on time and complete. OTIF = (430 ÷ 500) × 100 = 86 percent, below the typical 95 percent benchmark and worth escalating. Now the what-ifs. Breaking down the 70 failed orders: 45 were late but complete, and 25 were on-time but short-shipped. Fixing only the lateness issue (say, through better carrier selection) would improve OTIF to (430+45)/500 = 95 percent, hitting benchmark, while fixing only the completeness issue would improve it to (430+25)/500 = 91 percent, still below benchmark. This decomposition is critical for supplier conversations: "your OTIF is 86 percent" is far less actionable than "your OTIF is 86 percent, driven mostly by late delivery, not fill-rate gaps," which points the supplier at the specific process to fix. Next quarter, the same supplier improves lateness significantly (only 10 late orders now) but completeness issues persist (still 25 short-shipped): OTIF = (500-10-25)/500 = 93 percent, a meaningful improvement but still short of the 95 percent target, and the remaining gap is now concentrated entirely in fill-rate performance, which typically points to a warehouse-side or inventory-planning problem on the supplier's end rather than a logistics/carrier problem. Finally, compare line-level versus order-level measurement on the same data: if those 25 short-shipped orders were each missing just one line item out of an average 8 lines per order, line-level OTIF (evaluating each of the roughly 4,000 total lines individually) would show a different, typically lower, percentage than order-level OTIF, because line-level measurement penalizes every partial short-ship rather than only counting whole-order failures.
95 percent or higher is the standard target for top-tier suppliers in most retail categories. Below 90 percent typically signals a supplier relationship that needs active management. Below 85 percent, the cost of stockouts, expediting, and emergency replenishment becomes significant enough to warrant formal supplier escalation or dual-sourcing consideration.
Both views have value, but they answer different questions. Order-level OTIF is simpler to track and communicate. Line-level OTIF is stricter and usually a better leading indicator of underlying supplier process problems, since it catches partial short-ships that order-level measurement can mask. Track both if resources allow; if choosing one, line-level provides more diagnostic value.
Split every failure into "late but complete" versus "on-time but incomplete" versus "both late and incomplete." Each failure type points to a different root cause on the supplier's side (carrier/logistics issues versus warehouse/inventory issues versus systemic capacity problems), and a supplier scorecard conversation grounded in this breakdown is far more actionable than a single aggregate percentage.
Directly. A supplier's OTIF performance feeds the lead-time variability term in the safety stock formula; unreliable suppliers (low or volatile OTIF) require more safety stock to protect the same service level, which ties up more working capital. Improving supplier OTIF is one of the highest-leverage ways to reduce safety stock without accepting more stockout risk.
Trailing 90 days is a common standard for ongoing supplier scorecards, long enough to smooth single-incident noise but recent enough to reflect current performance. Trailing 12 months is more appropriate for annual sourcing and contract renewal decisions, since it captures seasonal patterns a 90-day window might miss.
Depends on the purpose. Simple order-count OTIF (as this calculator computes) is easiest to track and communicate broadly. Dollar-weighted OTIF gives more analytical weight to high-value orders, which is often more relevant for assessing actual business impact, since a failure on a $50,000 order matters more than a failure on a $500 order even though both count equally in a simple percentage.
This needs to be defined explicitly in the vendor agreement before measurement begins. Some retailers count an approved substitution as "in full" if quantity and category match; others count only exact-SKU fulfillment as in-full. Ambiguity here is a common source of disputes during supplier scorecard reviews, so the definition should be written into the contract, not inferred after the fact.
Two consecutive measurement periods below target (e.g., two consecutive months below 90 percent) is a common trigger for formal escalation. A single bad period is often noise; a sustained trend below target reflects a real operational problem worth addressing directly, potentially including dual-sourcing critical SKUs while the issue is resolved.
Deep-dive guides that explain the math behind this calculator.
The question is rarely which system to buy, it is which to buy first. A framework based on where the operational pain sits today.
Drop-shipping trades margin for reduced inventory risk. The trade prices out well for long-tail SKUs and poorly for fast-turning A-class inventory. The decision framework and where the hidden costs sit.
How lead time drives inventory arithmetically, the eight levers ranked by cash impact, when pure JIT fails, and the segmented hybrid model most sophisticated retailers converge on.
Set the trigger level that fires the next PO for an SKU. Reorder point combines expected demand during lead time with a buffer for the weeks that run hot. Get either half wrong and you either stock out or bury cash on the shelf. Outputs are the ROP itself, the lead-time demand behind it, the buffer as a percent of expected demand, and the days of supply that ROP represents at current sales velocity.
Open
Measure how many times a year your average inventory sells through and gets replaced. The single most consequential operational KPI in retail. It connects buying decisions, warehouse cash, markdown risk, and finance targets into one number. The turn ratio arrives converted into days and weeks of supply, together with the working capital a one-turn improvement would release.
Open
Handpicked benchmarks, templates and guides to help you dig deeper.
Five core calculators every buyer, merchandiser and category manager reads together. Open the metric that is behind, and let the others sanity-check it.
Percent of revenue kept after paying for the goods. The anchor number on the retail P&L.
Open calculator
Percent added on top of cost to reach the selling price. The buyer’s pricing language.
Open calculator
Gross profit per dollar of average inventory. The honesty check on margin and turnover.
Open calculator
Units sold as a percent of units received. The leading indicator for markdown timing.
Open calculator
How many times average inventory sells through in a year. The core inventory KPI.
Open calculator