How to choose a Salesforce partner, and how to tell before you sign
Credentials are the only thing that compares cleanly between firms, which is why every shortlist is built out of them. They are also close to the weakest signal available. This is how to get evidence instead.
Platform & ArchitectureEvery Salesforce partner shortlist gets built out of the same three numbers, because they are the only ones that compare cleanly across firms: certifications held, tier attained, case studies produced. They fit in a spreadsheet, they can be verified in an afternoon, and a procurement process can defend a decision made on them, which is an underrated virtue when that decision is later questioned.
They are also close to the weakest signal available, and the distance between what they measure and what you are actually buying is where most bad partner choices are made.
We should say where we stand before going further. We are a Salesforce consulting partner, we compete for the work this article is about, and you should discount the argument accordingly. The only test that matters for a piece like this is whether it stays useful to someone who reads it and then hires one of our competitors. Where a criterion below happens to favour a large global integrator over a firm our size, we have said so rather than quietly steering around it. Nothing here ends in a request to call us.
What you buy when you engage a partner is four or five named individuals, for a period of months, making decisions you will live with for years after they have moved to another account. Almost nothing on a credentials slide describes those people or those decisions.
What a certification actually certifies
A Salesforce certification is a statement about one person on one date: they sat an exam and passed it. That is real and it is not nothing. Somebody covered the syllabus and has enough platform vocabulary to be useful in a design conversation. The list of what it does not tell you is longer than the claim.
It does not tell you the certified person will be on your project. Certifications are counted at firm level and staffed at project level, and those two facts have no obligation to meet. A firm can hold two hundred certifications and put none of the holders on your account.
It does not tell you whether they have ever delivered anything. Exams test recall of what the platform can do. They do not test whether somebody can sit with a sales director who wants a field that will destroy the forecast and talk them out of it. Delivery is mostly that conversation, repeated.
It does not tell you the firm can deliver. A firm is a delivery method, a set of habits about estimating and testing, and a way of handling the moment a programme is late. Individual credentials are silent on all of it, and very good consultants with no shared method are a common and expensive thing to buy.
The useful version of the question is not how many certifications the firm holds. It is which certifications the four people assigned to your project hold, when they last did the work those certifications describe, and what they built most recently. That question has an answer, and the quality of the answer is informative by itself.
What AppExchange tells you, and what it does not
The AppExchange consultants directory is the closest thing to a neutral register, and the right place to start, because the information is held by Salesforce rather than written by the partner.
Competencies. Under the FY27 Salesforce Consulting Partner Program, partners earn Competencies at two levels, Accredited and Expert. Competencies replaced the FY26 Navigator distinctions. Previously earned Navigator distinctions were mapped across to their equivalent FY27 Competencies and are marked Legacy in the Competencies Dashboard, which is worth knowing because it explains why the language you meet in the market is inconsistent. Earned Competencies appear in the search filters on AppExchange, and partners receive a badge to use in their own marketing.
That filter is genuinely useful for building a longlist. It narrows the field to firms that have met a defined bar in a defined area, which is more than a website can tell you. What it does not do is rank the firms that clear it, and it says nothing about who inside the firm carries the capability.
Tier. Consulting partners advance from Provisional to Select, and Salesforce's FY27 consulting programme material also describes Summit as the top consulting tier. We were not able to open the current FY27 brochure directly, so treat that as the shape of the structure rather than a complete list, and confirm the current tier and Competency names with Salesforce before writing either into a scoring matrix. Programme structures change on Salesforce's fiscal calendar, not yours, and a procurement document naming a superseded distinction reads badly to the firms you most want bidding.
Reviews. AppExchange reviews are the most underused artefact in the process, and reading them properly takes twenty minutes. Ignore the star rating; it compresses everything interesting out. Read the text for three things: whether reviews describe named phases and named problems or only enthusiasm, whether the same consultants are named repeatedly across years, which is a retention signal no firm will volunteer, and the distribution over time, because a cluster inside one quarter then silence usually means a review campaign rather than a review pattern.
The badge on their own website is not evidence
Badges persist on websites. Nothing reaches into a partner's site and removes a graphic when standing lapses, changes or is superseded, and there is rarely bad intent involved: a marketing page gets built once, and nobody audits it annually. Much of the badge language still in circulation is FY26 Navigator language, which has been replaced. A site showing Navigator terminology is not necessarily misrepresenting anything. It is showing you a page nobody has updated.
So do not evaluate standing from a partner's own site, in either direction. Check the current position on AppExchange, where the record is held by Salesforce. If it differs from the website, ask without accusation. The answer is usually mundane, and how the firm handles a mundane discrepancy is itself a small piece of evidence.
Evaluation signals, and what each one actually proves
Most signals prove something narrower than the thing you want to know, and each has a specific question that closes the gap.
| Signal | What it actually proves | How it can be faked or inflated | The question that tests it |
|---|---|---|---|
| Certification count | Named individuals passed exams | Counted firm-wide, scales with headcount, none of them need be on your team | Which certifications do the four people assigned to us hold, and when did each last do that work? |
| Competency or tier | The organisation met a defined bar | Says nothing about your delivery team, and the language may be out of date | Verified on AppExchange rather than the website, and which named people carry it? |
| Badge on partner website | The badge was placed there at some point | Persists after it lapses, and may be superseded FY26 Navigator language | Show me the current standing on AppExchange |
| AppExchange reviews | Some clients were willing to be asked | Solicited at the happiest moment, campaigns cluster in one quarter | Which of these clients had a phase go badly, and may we speak to one? |
| Case study | A project happened and ended well enough to publish | Omits duration, overrun, scope cut, and who has since left | What was in the original scope that is not in this case study? |
| Senior consultant in the pitch | That person was available for a two-hour meeting | Presales staffing is not delivery staffing | What percentage of their time is contracted to us, and for which weeks? |
| Reference call | The partner has at least one client willing to speak | References are selected and briefed | May we speak to your two most recent clients, whoever they are? |
| Industry experience | Somebody at the firm has seen the vocabulary before | Vocabulary is not the regulatory or operational detail | Which specific constraint in our sector changed your design, and how? |
| Low price | The bid is low | Usually reflects a narrower reading of the problem | Which assumptions produced this number, and what happens if each is wrong? |
Nothing in the middle column is dishonest. It is what the signal genuinely means. The failure is treating it as a proxy for something larger.
The substitution problem, and the clause that fixes it
The most experienced people at a consultancy are in the pitch because winning work is what they are best at. That is a rational allocation of a scarce resource and every firm does it, ours included. The problem arrives at handover, when the architect who designed your solution moves to the next bid and the delivery team you get is one you have never met.
This is not usually a bait and switch. It is the ordinary consequence of a resourcing model where projects start and finish on unpredictable dates. The effect on you is identical either way, and the design intent that lived in the pitch architect's head does not transfer in a document. Only one thing reliably addresses that, and it is not a promise in a meeting. It is contract language.
Ask for a schedule that does four things. It names the individuals in the delivery team, with role and the specific weeks each is committed for. It states the percentage of their time you have contracted, because a named architect at ten percent is a reviewer rather than an architect, and that is worth settling before it becomes a grievance. It gives you the right to interview and reject a replacement, which costs a firm nothing if it intends to honour the staffing and constrains it usefully if it does not. And it requires notice of any change, with a handover period where both people are on the project rather than a clean swap on a Monday.
Expect a genuine limit. No firm can guarantee an individual against illness or resignation, and a partner who agrees to an absolute commitment with no exceptions is either not reading the clause or not intending to be held to it.
How to read a case study, which is mostly the omissions
A case study is a marketing artefact produced with the client's approval, which means two parties have already agreed what it will not say. Read it for the shape of the hole.
Start with duration. Most case studies give a start or a go-live, rarely both, and almost never the original planned date next to the actual one. Absent duration is the most common omission because it is the most common overrun.
Then the team. Ask which of those people still work there. Consulting has real turnover, and the capability that delivered a flagship project three years ago may have left since, taking the method with it.
Then the scope. Almost every published project ended smaller than it started, usually through good project management rather than failure: something got deferred so the rest could land. The deferred piece is invisible in the write-up, and it is often the hardest piece, which makes it exactly the part you want to hear about.
The question that turns a case study into evidence is short: what was in the original scope that is not in this case study, and why? A good answer names the deferred item and the reason without defensiveness. A bad answer is that everything in scope was delivered, which is true of almost no project of any size and mostly tells you the person answering was not there.
The questions that separate delivery from a pitch
A list of questions is useless without a sense of what a real answer sounds like, because every firm has heard the questions and the polished answer is the failure mode. Five that are hard to rehearse.
"Tell me about a project that went badly and what you changed afterwards." A good answer is specific and slightly uncomfortable: a named failure, an attribution that includes the firm's own contribution, and a concrete change to how they now work, which they can describe because they use it. A bad answer is a failure that is really a compliment, usually a client who would not take the advice, or a story with no change at the end of it.
"What would make you tell us not to do this?" A good answer arrives quickly and contains real conditions: no business owner, a data source nobody has profiled, a deadline the scope cannot fit. It means they have declined work before. A bad answer is that they would find a way to make it work, which sounds like commitment and is the absence of a threshold.
"Walk me through your last go-live, hour by hour." A good answer is boring and full of texture: the cutover window, who was on the bridge, what broke, how long the fix took, the hypercare rota in week one. It cannot be invented on the spot. A bad answer stays at the level of phases and gates and a checklist, which is what somebody says when they have read the method rather than run it.
"Which of our requirements do you think is wrong?" Ask this once they have seen your requirements. A good answer names one and explains the reasoning, and it may well be a requirement you are attached to. A bad answer is that the requirements look comprehensive. You are buying judgement, and this is the cheapest available test of whether there is any.
"What do you need from us, and what happens if we do not provide it?" A good answer is precise about your obligations: a decision maker available on named days, data extracts by a date, users released for testing in a named week, and the consequence when each slips. A bad answer is that they will work with whatever you have, which relocates every dependency risk into the invisible part of the plan.
What the commercial model does to behaviour
Every commercial model creates incentives, and the incentives predict behaviour far better than any statement of values. The question is not which model is best, because none of them is, but which distortion you would rather manage.
| Model | What it genuinely incentivises | Where it goes wrong | Suits |
|---|---|---|---|
| Fixed price | Tight scope control and efficient delivery | Change requests become the margin, and the answer to a good idea in week six is a form | Well-defined work with stable requirements and a clear boundary |
| Time and materials | Flexibility, and a willingness to follow the problem | No structural pressure to finish, because duration is the revenue | Discovery, exploratory work, estates nobody fully understands yet |
| Outcome-based | Focus on a measurable result | Defining the measure becomes the negotiation, and the measure is rarely the outcome | Cases where the metric is unambiguous, agreed and already instrumented |
Fixed price does not mean the price is fixed. It means the scope is, and that discipline is real: both sides have to be specific up front. The failure mode is that in month three you learn something that should change the design, and the model has priced that learning as a variation.
Time and materials is honest about uncertainty and puts the estimating risk on you. It fails quietly, because nothing in the structure creates the moment where somebody has to say the work is finished.
Outcome-based sounds like the answer and is the hardest to write. The measure has to be defined before either party knows enough to define it well, and once written, both sides optimise the measure rather than the outcome. If you cannot state exactly how the number will be calculated, from which system, by whom, and what happens if the baseline is disputed, the model is not ready.
Do not pick a model because it appears to transfer risk. Risk is not transferred by a contract, only priced by one.
What happens when the engagement ends
A partner who has made themselves necessary is a partner you chose badly, and the time to test for it is before signing, because afterwards you have no leverage.
Ask what handover means concretely. Not whether documentation is provided, which everybody says yes to, but what documentation, in what form, and where it will live. Design decisions and their reasoning are the part that matters, because the next administrator needs to know why the automation looks like that, not what it does, which they can read.
Ask where the code and configuration live. If the repository is the partner's and access ends with the engagement, that is a dependency you bought without noticing. Ask whether your own administrator can sit inside the delivery team during the build, because knowledge transfer scheduled as a phase at the end is mostly a slide deck. Then ask what the first ninety days after handover cost. A partner with a real answer describes a tapering arrangement and quotes it.
The prior question, worth settling before you scope the engagement at all, is whether you should build an internal capability instead of a permanent external one. That trade is worked through in managed services or an in-house admin: how to decide honestly, and the answer is genuinely situational.
Run a paid discovery, or a small first piece of work
This is the most effective evaluation technique available and it is not close.
Every method above predicts how a firm will behave from artefacts produced by people whose job is to win the work. A paid discovery, or a small and genuinely useful piece of delivery, replaces prediction with observation. You watch the assigned team work on your estate, with your stakeholders in the room, for two to six weeks. You learn how they write, how they estimate, whether they surface bad news early or late, what they do when one stakeholder contradicts another, and whether the senior person from the pitch appears at all.
Pay for it. Free discovery is a sales activity, staffed accordingly, and scoped to produce a proposal. Paid discovery is delivery, staffed accordingly, and scoped to produce something you own. Insist that the output is yours to take to another firm, and say so up front. A partner confident in its work will agree, because a good discovery makes them the obvious choice anyway.
Two structures work: a discovery producing an architecture and a plan you can hand to anyone, or a small self-contained delivery with a real go-live, one integration, one process, one team. Run it with two firms if the programme is large enough to justify the cost. That is more expensive than a beauty parade, and dramatically cheaper than choosing wrong on a programme that runs for a year.
Your estate should change your shortlist
The right partner is a function of the estate, and firms of different shapes are genuinely better at different things. This is where we say things that do not favour a firm our size.
Regulated organisations. If you sit under a regulator, in financial services or health, or you handle data with residency constraints, a large global systems integrator often is the better answer, and not because the consultants are better. The firm already holds the certifications your risk function will ask for, has been through the audits, carries the insurance limits your legal team will insist on, and answers a security questionnaire from a library rather than from scratch. A smaller firm can do excellent work and still fail your onboarding process before anybody assesses the work.
Multi-country rollouts. Multiple time zones, local legal and tax variation, and a template that has to survive contact with a country that does business differently. This needs a firm with people in those regions and a rollout method it has used before. It is a genuine scale advantage, and a boutique will usually struggle with it.
A fifty-user org. The opposite. A large integrator will assign a competent but junior team, run a delivery method built for programmes ten times the size, and charge you for governance you will never meet. A boutique or an experienced independent will do it faster and cheaper. The risk you take is concentration: one or two people, no bench behind them.
A complex estate with one demanding product. Revenue Cloud, CPQ, Marketing Cloud, Data Cloud. Product depth beats general scale, and product depth lives in individuals rather than firms, which returns you to naming the team.
An org carrying accumulated debt. If the org is eight years old and nobody can safely change anything, you need diagnosis before delivery. That is a different skill and often a different, smaller supplier. Do the first pass yourself with a technical debt audit you can run on your own org: arriving at a partner conversation holding your own findings changes the conversation considerably.
The cheapest bid, and the most expensive one
The cheapest bid is usually the one that has understood the least.
This is rarely dishonesty. A firm that has not examined your integration, not asked about data quality in the source system, and not discovered that two departments define a customer differently produces a smaller number, because it is pricing a smaller problem. The bid is an accurate price for the work as understood, and the understanding is what is deficient. You pay the difference later, as change requests, at a rate set when you have no alternative.
The reverse claim is not true. The most expensive bid is not therefore the most thorough. It may be a firm with a higher cost base, a bid loaded with governance roles you will never meet, or a firm that does not much want the work. Price does not rank quality in either direction.
What is diagnostic is not the number but the assumptions underneath it. Ask every bidder to list the assumptions their price depends on, and what happens if each is wrong. The firm that has understood the problem produces a short, specific, slightly alarming list. The firm that has not produces generalities. Then compare across bidders: an assumption only one firm made is often the most interesting thing in the process, because one of them saw something the others did not.
The mistakes most likely in the first fortnight of any engagement are set out in Salesforce implementation mistakes, ranked by what they cost to undo. A partner who names the top three unprompted has been there.
The case for not using a partner at all
This is real, it applies more often than the consulting market admits, and no partner will raise it with you.
If your org is small, running standard functionality, with a competent certified administrator, no integrations more complicated than a well-supported connector, and no regulatory exposure, you probably do not need an implementation partner. You need that administrator, a Trailhead habit, the discipline to change one thing at a time, and a specialist you can call by the hour for the two or three decisions that are genuinely irreversible.
That works because Salesforce is unusually forgiving of iteration. The exceptions are concentrated in a small number of early data model choices, which is exactly where a few hours of paid specialist advice beats a full engagement. The data model decisions you cannot cheaply undo covers which ones they are, and they make the right agenda for that conversation.
The case for not using a partner weakens on specific triggers rather than on general size:
- Integration to a system of record, where the failure mode is silent data corruption rather than an error message.
- A migration carrying history that cannot be reconstructed if it is loaded wrongly.
- Multi-country, multi-entity or multi-currency requirements.
- A regulated process where being wrong is a reportable event rather than an inconvenience.
- A deadline that makes learning on the job an unacceptable risk.
- An internal team with no capacity, which is different from no capability, and is the most common trigger of all.
None of those is about org size. Two hundred people with one integration to a finance system is a harder problem than two thousand running standard Sales Cloud.
The honest middle position, under-sold because it is worth less to firms like ours, is a partner for the irreversible parts and your own people for the rest. Buy the architecture, the data model, the integration design and the migration. Build the layouts, the reports, the automation and the ongoing change internally. It costs less, it leaves capability behind, and it produces an org somebody in the building understands.
The decision you actually have to make
Everything above reduces to one choice, and it is not which firm to hire.
It is how much evidence you are willing to buy before you commit. A traditional selection, run on credentials, references and a proposal, is cheap and fast, and it decides on artefacts written by people whose job was to produce them. A selection that includes a paid discovery, named individuals in a contract, references you chose yourself, and a handover test written as an acceptance criterion, costs more up front and takes longer to reach a signature.
The second is not always right. On a small, well-understood piece of work it is disproportionate, and the time spent evaluating exceeds the cost of choosing wrong. On a programme that will run for a year and shape how the organisation works for five, it is the only version that makes sense, and the extra cost is a rounding error against the difference between a good outcome and a mediocre one.
So the question to settle first, before you write a single evaluation criterion, is which of those two situations you are in.



