How to evaluate a productivity tool before you rearrange your work around it
Six questions in a fixed order, and a two week trial built to answer them. The order does more work than the list does.
Almost every bad software adoption we have watched had the same shape. The tool was good. The evaluation started in the wrong place.
Start with the job, not the tool
The usual sequence is that you encounter a tool, it looks impressive, and you begin looking for somewhere in your work to put it. That is backwards, and it is backwards in a way that is hard to notice from the inside, because the search for a use is genuinely enjoyable and produces a lot of activity.
The alternative is to start from a job that is currently going badly and describe it before you look at anything. Not “I need a better note app”, which is already a solution, but the actual sequence: where the work starts, what you do with it, where it ends up, and which specific step you would pay to remove. Written down, that is usually three or four sentences, and it does two useful things. It gives you a test the tool either passes or does not, and it makes it obvious when you are about to adopt something that solves a different problem well.
This is also the point at which a fair number of evaluations should stop. If you cannot name the step you want removed, the honest conclusion is not that you need a better tool. It is that you do not yet know what is slow, and the next move is to watch your own work for a week rather than to install anything.
The six questions, in order
These are the same six questions every write-up on this site is built around, and the order is not cosmetic. Each one is only worth asking once the previous one has an answer, because a good answer to a later question cannot rescue a bad answer to an earlier one.
-
What does it actually do
Stated plainly, in your own words, without any of the vendor's adjectives. If you cannot write one sentence that would let a colleague picture the thing, you do not understand it well enough to evaluate it, and the marketing has done its job.
-
Where does it sit next to what you already use
The important question is not what it integrates with. It is what it replaces. A tool that adds a step to an existing sequence has to be enormously good to be worth it; a tool that removes one only has to be adequate.
-
What does it cost to get to the first useful minute
Accounts, permissions, imports, configuration, a settings pass, teaching it your vocabulary. This is the number that predicts abandonment better than any other, and it is never on the pricing page.
-
What does it cost against what it saves
Including the free tier and, more importantly, where the free tier stops. The interesting question about a free tier is not what it gives you, it is what usage pattern pushes you off it, because that is the price you are actually being offered.
-
What does it not do
Every tool is bad at something, and if you cannot name the thing after a fortnight of use you have not used it hard enough to know. Finding the limitation is the point of the trial, not a disappointing side effect of it.
-
Who is it wrong for
Then ask honestly whether that description fits you. This is the question that most often changes the decision, and it is the one people skip, because by this stage in an evaluation almost everybody wants the answer to be yes.
The trial: two weeks, two jobs
A trial is worth running properly or not at all. The design below is deliberately awkward in one respect, which we will come to.
The job you chose it for
Use it only for the specific job you wrote down at the start. Nothing else. You are testing a claim, and broadening the test is how you lose the ability to answer it.
A job you did not choose it for
Now use it for something adjacent. This is where you find out whether it is a tool or a feature, and whether the second use costs you anything in the first.
The awkward part is week two. The instinct after a good week one is to commit, and the instinct after a bad week one is to stop. Both are premature. A tool that only works for the exact job you bought it for is a fine outcome, but it is a different outcome from a tool you can build on, and you should know which one you have before you rearrange anything around it.
Two rules make the fortnight worth more than it costs. First, do the setup properly on day one, including the settings pass; a tool evaluated on its defaults is often a different product from the same tool configured. Second, keep a running note of every friction, one line each. Not a report. A list. At the end of two weeks the list is the finding, and it will contain things you would otherwise have forgotten within a day of feeling them.
Six ways a trial misleads you
Trials systematically overstate how well a tool will work, and they do it in six predictable ways. Knowing the shape of the distortion is most of the defence.
The enthusiast fortnight
You are paying attention in a way you never will again. Effort you are spending without noticing gets attributed to the tool as though it were free.
Clean starting data
A new workspace is small and tidy. Almost every tool in this category is good at fifty items and revealing at five thousand, and a fortnight will not get you there.
The easy sample
Without meaning to, you feed it the work it is likely to handle well. The hard cases quietly go to your old method, and you never notice you routed around the problem.
Novelty as evidence
New tools are enjoyable, and enjoyment reads as productivity from the inside. The honest version of this feeling is that you have been more engaged, which is real but does not survive contact with week six.
No maintenance yet
Nothing has broken, updated, changed a permission, or needed re-authorising. In two weeks you have seen the tool in perfect health and nothing else.
You are alone in it
If the tool will eventually involve colleagues, a solo trial has tested the easiest version of the problem. Shared use is a different product with the same name.
The correction for all six is the same and it is unglamorous: assume the tool will be somewhat worse than the trial suggested, and ask whether it would still be worth adopting at that lower level. If the answer is yes, you have a real finding. If the answer depends on the trial being representative, you do not.
What “it didn't stick” usually means
When a tool is abandoned, the reason given is almost always that it was not good enough. That is occasionally true. More often the tool was fine and one of three other things happened.
- It required a decision every time. Anything you have to remember to use loses to anything already in the path. This is the single most common cause of abandonment and it has nothing to do with quality.
- It sat next to the old thing instead of replacing it. Running two systems is more work than running either one, so the cheaper of the two wins, and the cheaper one is always the one you already know.
- The first useful minute was too far away. If a tool needs configuration before it repays anything, it is competing against your motivation on the specific afternoon you installed it, and that is a weak opponent to be relying on.
All three are predictable from questions two and three, which is the argument for asking them in order. None of them are visible on a feature comparison.
Write the decision down
Whatever you conclude, write four lines somewhere you will find them again: what you tried, what you were trying to fix, what you decided, and what would change your mind. The last line is the one that pays.
Tools change. A limitation that ruled something out this year is a release note next year, and without a record you will either re-evaluate it from scratch or, more likely, carry a stale opinion for years. “No, because there is no offline mode” is a decision you can revisit in ten seconds when offline mode ships. “No, we looked at that one” is not.
The whole method, in five lines
- Describe the job before you look at any tool.
- Ask the six questions in order and stop early if an early one fails.
- Two weeks: the job you chose it for, then one you did not.
- Discount the trial, because trials flatter.
- Record the decision and the thing that would reverse it.
None of this is difficult, and that is rather the point. The reason tool decisions go badly is not usually a lack of analysis. It is that the evaluation started with a product instead of a problem, and everything after that inherited the mistake.