← Field Notes

Velocity is an alignment problem.
Alignment is an information problem.

Your team is built of people who believe in a shared mission, staffed with people who disagree on how to achieve it. This tension is good, but then why does your team feel misaligned? The reason for this feeling is because everyone, including your AI assistants, is reasoning from different ground truth.

Why it matters: With AI, information is cheap to create, and with more information, misalignment is more common. Every product builder has their own ground truth of screenshots, anecdotes, and AI chat history. Each of these leads to a different reality for each product builder, which compounds into misalignments that cause disjointed conversations, feature fragmentation, and product degradation. Fortunately this is not a culture problem, it's an infrastructure problem.

A shared, validated corpus.

The solution is a shared corpus of validated insights, constantly updated with new information. With this shared corpus, continually being enriched by additional data, everyone – and everything – is reasoning from a single source of truth.

The atomic unit of insight: Flash Findings.

One page or less:

Decks are where insights go to die – we at Sibling Systems don't believe our work is born just to die. We want our insights to live, breathe, and have an impact on the world: our Flash Findings are not a passive report. Flash Findings are an active unit of validated knowledge, something small enough to digest and understand in two minutes, and structured enough for your teammates – meat based and digital – to make effective use of.

And the last bit there matters more and more as your team uses AI tools to draft PRDs, propose roadmaps, and spelunk in support data: having clear concise information is key. Your AI will reason over whatever corpus you give it to work with. While lots of data is great for pre-training models when you are actually using an AI model, concise input-output chains serve your business far better. Context degrades as the volume of mid-context material increases. Just dumping a ton of data is actually doing more harm than good, especially if it's just ten page slop docs. So: focused validated user findings, outputs that serve as inputs to your humans and machines, and creating the alignment that we so badly need.

The Impact Rubric.

We have implemented an "Impact Rubric" to build consensus about validity, what makes it into the corpus, and how to measure different types of evidence. Finding validity is not determined by the researcher, but by the team doing the work. We set up a base rubric including Confidence, Cross-Functional Impacts, User Impacts, Effort, and other measures that help your product builders know what matters.

Yes, but — validity by vote, at scale?

Yes, but doesn't that make validity a team-localized vote? How does that work for an organization at scale?

Well, we vote on whether a finding is decision-relevant, not globally true. This ensures a clear rubric and alignment on where the inputs can work. This also, very importantly, tells us when insights are expanding past their valid scope and need to be reevaluated. With these measures, we begin to see an emergent decision making structure based on aligned understanding of insights, weights, and impacts. Localized truths first, additional information collected to recalibrate as the terrain expands.

The bottom line: Velocity without shared alignment is ten people sprinting in ten different directions. Alignment does not mean agreement either – it just means we are all discussing the same base facts.

And Rolling Research is a great way to get the base facts, especially on highly contentious product decisions with fragmented information.

This week we have Hypothesis Hydration: write five hypotheses you have about a core feature flow, or a UI issue where user confusion is clear.

Quick If... Then... Because... hypothesis sample:

  • IF we create a "jackpot" feature to showcase paid products prominently
  • THEN users' trust will decrease in our product list
  • BECAUSE users will not differentiate paid products from highly rated/ filtered products

Next week we will show how five teammates and thirty minutes set you up to run a two-week scoped study that tests your hypotheses.

See also: Daring Fireball — The Load-Bearing Vocabulary of Claude