Running the method on myself: choosing the next hard problem

The setup

Circular is a practice: I take one hard B2B sales or RevOps problem per cycle, on a success-fee basis, and work it until it's done. After two paid engagements, I needed to choose the next one. Choosing which problem – and which person to work it with – is itself the hardest decision in the business, so I decided to run my own method on that decision instead of picking by gut.

The mechanism is a secretary-style rule. I run a warm outreach campaign across my own network, hold a structured first conversation with each person who's interested, and score every conversation on the same rubric. The rule is deliberate: don't set the bar from the first few calls, keep sampling through a calibration window of at least nineteen conversations, and only let a later conversation win by beating every prior one and clearing a hard fit test. The goal is one collaborator, or none. If no one clears the bar, I keep going rather than settle.

This is not a staged case study with a tidy outcome. It's an in-flight engagement, and I'm writing it while it's running. Five of nineteen first conversations are done. The bar is not calibrated yet. There is no closed outcome, and I'm not going to invent one.

The visible symptom

Looked at from the outside, this is an outreach campaign. The visible work is what you'd expect: reactivate dormant warm threads, get replies, keep a review queue healthy, move drafts through a review-and-send flow. If you measured the campaign the way most people measure outreach – messages out, replies in, queue depth – you'd be measuring the wrong thing.

The real problem isn't "send more DMs." It's to identify the single highest-leverage hard problem and the right counterpart to work it with, under a mechanism strict enough that I can't fool myself. That's a much harder thing than generating conversations. It means designing the first call to avoid anchoring on whoever's pain is loudest, and building verification into the conversation, because my working thesis is that a lot of problems people describe as "solved" are actually judgment or adoption gaps wearing a solved costume. The outreach is the surface. The selection discipline is the actual work.

The judgment moments

Every one of these is a call someone could reasonably disagree with. That's the point.

Warm reactivation over cold expansion, for this phase. I built the campaign to work only through already-connected threads. I have a cold, empty-connection-request sequence fully specified – and I deliberately haven't built it. The bet is that at the selection stage, signal quality matters more than reach: a lukewarm conversation with someone who already knows me is worth more than a cold one with a stranger, because I'm calibrating judgment fit, not filling a funnel. The disagreement is obvious – cold outreach would widen the sample faster. I chose depth of signal over breadth of sample.

The stopping rule is a hard constraint, not a guideline. No bar-setting until nineteen conversations. That means when a genuinely strong person shows up early – and one did – I am not allowed to stop and sign them. This is the single most counterintuitive discipline in the whole thing: a front-runner is not a stop signal. Every instinct says "finally, someone who gets it, close it." The rule says keep sampling, because "better than everyone so far" is meaningless at conversation five. Holding to that when an exciting prospect is in front of you is the hard part, and it's exactly where most people's selection process quietly breaks.

Stress-testing a "solved" claim into a disqualification. One conversation was with an agency operator who had, impressively, encoded his own deal-qualification judgment into a set of deterministic tools. On the surface: problem solved, strong peer, genuine rapport. I ran the verification probe anyway, and the conclusion flipped. The deterministic playbook exists, but the real gap is adoption and consistency: people on his team aren't using the tools enough. That's a real problem, but it's off my wedge – it's a change-management problem, not a hard data or systems one. So I disqualified a high-person-fit prospect on problem-fit grounds. The easy move was to talk myself into it. The probe is what stopped me.

A provisional hold instead of an immediate close. The strongest on-wedge signal in the set so far came from an operator working a genuinely hard graph and orchestration problem – mapping the dark gap between first touch and real buying interest across a buying-center. On problem shape and person fit, the best conversation yet. And I still didn't advance it. The outcome he described was directional, not sized, and the urgency was weak. So the call was "hold, provisional," pending a follow-up that puts numbers on it. Ranking a directional-but-exciting narrative below a strict evidence bar is uncomfortable when the alternative is momentum.

Personalization as an extension of the analyzer, not a separate system. When I built the thread-personalization piece, the tempting design was a standalone warm-reactivation system. I chose the smaller move: extend the existing conversation analyzer so a single analysis pass can also produce a draft reply, and only for threads that actually qualify for a rejoin – review-first, never auto-sent. Don't stand up a second system where an extension of the first will do, and keep a human in the loop on anything that goes out under my name.

The pattern, applied to myself: the whole campaign is a machine for not believing my own first impression. Each judgment moment – the warm-only scope, the stopping rule, the disqualification despite rapport, the hold despite excitement – is a refusal to stop searching too early.

What I built

Three lanes: outbound warm-reactivation drafts generated from thread analysis; an inbound lane for threads needing a reply; and a freshness sync that keeps thread state current. Drafts land as pending-review rows and never send themselves. Candidate filtering excludes group chats, threads where I sent last, and anything more recent than a four-week floor. Replenishment is bounded – a shallow target of pending drafts, a per-run analysis cap – so the queue stays reviewable by one person. The operator flow is three commands: analyze, queue, replenish. That's the labor. It exists to serve the selection discipline, not the other way around.

The numbers

There are no outcome numbers here, and that's the honest state of it.

  • 5 of 19 first conversations complete. The bar is explicitly not calibrated yet.
  • The replenishment queue targets 8 pending-review drafts at a time, with a cap of 16 analyzed per run. These are throttle settings that keep the queue human-reviewable; they are not performance metrics.
  • No reply rate, acceptance rate, conversion rate, or finalist count exists – and I'm not going to manufacture one for a case study.

The mechanism's own commitment – how far past nineteen I'll go – is simple: if no one clears the bar, I keep going. Publishing a precise number before the search is over would be exactly the kind of thing this case study is supposed to be against.

What I'd do differently

Force falsifiable quantification earlier in the first call, especially on urgency. The provisional-hold conversation would have resolved faster if I'd pushed for a number in the room instead of accepting a directional narrative and deferring the sizing.

Publish operating telemetry for the campaign itself. The ranking right now is mostly qualitative. Booked, replied, converted, queue depth over time – the campaign should generate the same legible operating data I'd insist on for a client. It's a little embarrassing that the meta-engagement is thinner on telemetry than the paid ones.

Keep an anti-bias watch on the source mix. Several disqualifications clustered around one profile: tech-native operators who build their own tooling. That's a sourcing artifact of where warm threads come from, not a verdict that the market is empty. Treat an off-wedge cluster as a finding about my network, not a truth about the world.

What the conversations keep proving

Self-referential, so there's no external customer to quote. The closest thing is the commitment the mechanism makes to itself: the rule only earns trust if it can cost me the person I most want to work with. If the discipline can't survive meeting someone great at conversation five, it isn't discipline.

The conversations keep proving the method's own thesis – that "solved" is usually a judgment or adoption gap in disguise. An agency operator, describing his deterministic tooling: "the gap I'm seeing is consistency in use – they're not using it enough." A founder who keeps cancelling planned hires: each model improvement "saves us a hire." Two people, two reasons the hard problem is somewhere other than where it first appeared. That's the whole reason to run the search instead of trusting the first strong call.

If your pipeline has a version of this problem, that's the conversation I want to have.