← BACK TO INSIGHTS
CUSTOMER RESEARCH METHODS9 MIN READ

Freight Carrier Performance Scoring Using Historical Load Data

How to build a carrier scoring system from load history

S
Grant Okonkwo
SEPTEMBER 17, 2026
CUSTOMER RESEARCH METHODS
SANSA LAB NOTES · 01

Freight in 2026 doesn't look like a rebound year, and anyone paying attention already priced that in. FTR Transportation Intelligence projects demand roughly flat against 2025, and the broader industry consensus points to a transition year defined by stabilization, not expansion. Capacity is shrinking too, down 3 to 5% as the market continues to contract. National truckload rates are around $2.24 a mile as of February 2026, flat enough that rate has stopped being the thing separating a good broker from a mediocre one.

What separates them now is knowing which carriers actually perform, on which lanes, and routing freight accordingly instead of re-negotiating the same rate for the tenth time. That knowledge already exists. It sits in the TMS, in old invoices, in a thousand email threads nobody's read twice. Data scarcity was never the problem, and pretending otherwise is how a lot of freight companies talk themselves out of doing the actual work. Almost nobody has turned what they already have into a scoring system that flags a bad carrier before it fails them again.

What historical load data contains and why most companies aren't using it

Every load a freight company has moved left a paper trail. On-time performance, tender acceptance, claims history, billing errors, dwell time at the dock, how fast someone answered a check call: it's all in the TMS or buried in an inbox, tagged to a specific carrier on a specific lane. None of it requires new instrumentation. It just sits there, unstructured, waiting for someone to connect the dots.

When most dispatch teams are asked how a carrier performs, the answer sounds like a bar review. "They're usually fine on the Asia-Europe lane." Somebody kept a spreadsheet once, updated it for a few months, then life happened and it stopped getting touched. That's institutional memory wearing a spreadsheet costume, and it decays the moment the person who built it changes jobs.

Running on vibes instead of data has a specific, repeatable cost. A carrier can slide for weeks, missing pickup windows a little more each time, before anything forces the issue. Contract renewals get decided on relationship and gut feel rather than a track record anyone bothered to pull. Good carriers don't get rewarded with more freight, and bad ones don't get corrected until a shipment blows up in front of a customer. Descartes' whitepaper on carrier analytics puts a number on what fixing this is worth: analytics applied across carrier rationalization, negotiation, and strategy can cut overall freight costs by 15 to 25%. Not by finding cheaper rates, but by routing to carriers whose actual, total cost of service is lower. That's a different exercise entirely from shopping for a discount, and most people still confuse the two.

The data needed to score carriers objectively already sits inside any company that moves freight. The gap is structure, full stop.

The six performance signals that matter, weighted by lane

Six signals do most of the work, and all six come straight out of load history that's already being generated.

On-time pickup and delivery ties most directly to what a customer experiences. Tender acceptance rate shows whether a carrier reliably says yes to loads on the lanes it's contracted for, and a carrier that's flaky on a critical lane is a risk sitting on the books rather than a mere inconvenience. Claims ratio catches the carrier who looks cheap on rate and turns out expensive once damage and loss get counted. Incorrect freight invoices are a well-documented industry problem, and bad accessorial charges are a well-known driver of those errors. That's a leak, plain and steady. Dwell and detention time drags on top of it: Driver detention imposes billions of dollars a year in cost to the trucking industry, and the carriers contributing most to that number show up in load history for anyone who bothers to look. Responsiveness rounds it out, invisible right up until a shipment goes sideways and nobody answers the phone.

Goodship's common starting weighting runs on-time delivery at 35%, tender acceptance at 30%, claim rate at 20%, and committed volume capture at 15%. That's a fine default and a bad universal rule, and treating it as gospel is where most scorecards go wrong. A temperature-controlled lane where a late delivery means a truckload of spoiled produce should weight on-time delivery far more heavily than a bulk commodity lane where the real risk is a carrier ghosting the tender.

Getting the weights right means sitting down with the carrier-facing team, the people fielding the angry calls, and asking what actually hurts. They'll say it straight, and their buy-in decides whether a scorecard gets used or gets ignored by week three. Segmentation matters just as much: a carrier excellent on dry van regional freight can be a disaster on refrigerated long-haul, and an aggregate score smooths right over that gap. Score by lane. Score by service type. Anything else is an average pretending to be an insight.

How AI carrier scoring differs from a spreadsheet scorecard

A spreadsheet scorecard has a shelf life measured in the time it takes to compile it. By the time a PDF summarizing last quarter lands in someone's inbox, the market has already moved, and the conversation about a specific carrier has to start over from a stale baseline. Carrier behavior doesn't wait politely for the report to catch up, and treating a quarterly review as real-time intelligence is the mistake most ops teams don't realize they're making until a carrier they trusted craters mid-contract.

AI-driven scoring runs on a different clock. It works across a substantial window of load history to catch recurring patterns and seasonal swings, layered against live market signals, and it updates continuously instead of quarterly. OS For Your Business reported a dynamic scoring approach in 2026 that tracks a wide range of performance metrics per load, and reported a 35% drop in service failures along with a 20% lift in customer satisfaction scores. On the matching side, Groovy Web reported that AI load matching, modeling lane preference, equipment availability, GPS position, hours-of-service compliance, rate history, and performance record all at once, hits first-match placement rates of 94 to 98%, against roughly 67% for manual brokerage.

The real shift sits in what the score gets used for. It stops living in a report someone opens once a quarter, because the system acts on it directly. Volume shifts automatically toward carriers with strong on-time records, high first-attempt success, and clean settlement history, and away from carriers trending in the wrong direction, without a human having to spot the trend first. For a mid-size brokerage juggling relationships across hundreds or thousands of carriers, that's the actual job the AI does: narrow the field to the 20 or 30 genuinely good fits for a given load, fast and consistently, instead of leaning on whichever broker happens to remember a name.

What enterprise deployments have already demonstrated about automated scoring

None of this is hypothetical. C.H. Robinson reports more than 30 AI agents running across its global network, automating over 3 million tasks. Its Orders Agent reads a tendered email, attachments included, and builds a complete order in about 90 seconds, processing roughly 5,500 truckload orders a day and saving an estimated 600 hours of labor. The Quoting Agent returns a customer-specific price quote in 32 seconds and has processed over a million of them. The LTL freight classification agent, once trained, classifies freight in a matter of seconds, a dramatic reduction from the time a manual classification process typically requires.

Uber Freight isn't far behind, with AI-powered platforms that have demonstrably reduced empty trips and improved matching across its network. Alvys, a freight software provider, launched an agentic platform called Alvys Foundry with more than 20 pre-built agent templates covering detention processing, document handling, rate audits, tracking, compliance, and claims, with the option to build custom agents or work with Alvys' own engineers, Artificial Intelligence News reported. Trimble unveiled an AI-powered TMS in November 2025 built around seven modules covering order acceptance, load building, planning, fleet readiness, and live tracking.

Every one of those is an in-house build at a company large enough to fund its own engineering org. The logic scales down fine to a smaller shop. The payroll required to build it from scratch does not, and that gap is exactly where most mid-market freight companies stall out and give up on the idea.

Building a scoring system without an in-house engineering team

A 50-to-100-person brokerage doesn't have a bench of machine learning engineers sitting around, and pretending otherwise leads to expensive, half-finished projects. Internal AI builds in logistics are widely known to stretch well past a year before reaching production, which makes the enterprise path a non-starter for anyone without enterprise headcount.

The alternative is embedding capability instead of building it from zero. Debales AI says a full-stack platform approach can reach deployment in 4 to 8 weeks, mostly by skipping the custom API work that eats most of a build's timeline. The model that fits best here is an embedded practitioner who works directly inside the freight company's actual data, its TMS records, its invoices, its carrier email threads, rather than a vendor who ships a generic dashboard and disappears.

That kind of engagement runs in three steps. Diagnose the actual bottleneck before building anything: for most freight operations, that bottleneck is carrier selection or load allocation, because the data there is richest and a bad decision there raises costs fastest. Build a scoring system shaped around the company's real lane mix and carrier base. Then let it compound, since the model gets sharper as it accumulates more load history, and the first 90 days of live data usually beat any static scorecard by a wide margin. One freight company recovered more than 160 hours a month by automating freight rate comparisons across roughly 300 routes tracked daily, work that used to eat an analyst's whole week.

The upside runs past efficiency for its own sake. A five-person dispatch team that used to route freight on gut feel can start routing on performance data instead, take on lanes it would have turned down before, and make carrier calls that used to require pulling in a specialist.

Turning carrier scores into routing decisions that compound over time

A scorecard that appears only at the quarterly review is a nice-to-have. A scorecard that feeds the dispatch board every morning changes how the business runs. Acting on the number before the quarter ends is what changes the business, not sophistication, and companies that treat scoring as a reporting exercise rather than a routing trigger are wasting the data they already paid to collect.

Routing logic built on real scores looks like this: carriers with strong delivery performance against schedule, clean claims history, and high tender acceptance get first look at contracted lanes. Underperformers get pushed down the list before they cause a failure, not after. That same lane-level data changes the negotiating table too. Contract renewals stop being a conversation about goodwill and start being a conversation backed by a paper trail, with specifics on where a carrier is strong and where volume moves if service doesn't improve.

The compounding part is mechanical and unglamorous. Every load scored adds to a carrier's lane-specific history, the model's predictions sharpen a little more, routing improves a little more, service failures decline, and margin holds without anyone adding headcount to protect it. Forwarders running data-driven carrier allocation and negotiation typically see on-time delivery improve by 12%, not from squeezing a rate but from routing freight to carriers whose true cost of service is lower.

What it removes is the ceiling: the one that forces a broker to hire another dispatcher just to keep up, or eat a claim it should have seen coming, or turn down volume on a lane it can't confidently staff. None of this needs a finished model on day one. It needs the load history already sitting in the TMS connected to a scoring framework, run for 30 days, so the first lane-level results show exactly where freight has been misrouted all along. That first result is usually enough to pay for the next improvement.

Want results like these for your business?

Book a free 30-minute consultation. A focused conversation, no commitment.

Talk to Us
RELATED INSIGHTS