Your Local File Should Not Have to Argue With Your Database

Sync bugs usually all start the same way. Two copies of something, both of them mostly right, and no written rule about which one wins.

The problems occur when you don’t notice The bug. When the file says one thing and the database says another. It’s not a problem until it is. And then you have to spend time figuring the why and when’s of the drift.

So let’s talk about authority. Not storage, not sync, not “where does the data live.” Authority. Which copy is allowed to be right when two copies disagree.

The Question Nobody Writes Down

Most systems that hold the same data in two places never actually decide this. The decision gets made accidentally, by whichever code path happened to run last, and then it gets re-made differently by the next feature.

The failure is duplication without a stated rule.

Here’s a concrete version. I run a content pipeline for this blog. Posts are Markdown files with YAML frontmatter sitting in a directory. There’s also a Turso database holding metadata about those same posts. Two copies of what looks like the same information.

Ask the naive question, “which one is the source of truth,” and you get a bad answer, because the honest answer is neither, and both, depending on the field.

Split Authority by Field, Not by Store

You might try to pick an authoritative source based on store. Files win, or the database wins. But the useful granularity is usually the field.

In my pipeline it breaks down like this:

  • Post content and tags: the Markdown file wins. The frontmatter is authoritative. If the database has a different tag list, the database is wrong, and it gets rebuilt from the file.
  • Scheduling: the database wins. What time a post goes out, what slot it holds, whether it’s been claimed. The file does not get a vote.

Those are different answers for the same post, and that’s fine, because each one is written down and each one has a reason.

The content lives in the file because content is the thing I edit by hand, in an editor, with Git history behind it. I want git log to be the real record of what changed. Putting that in a database would mean my writing history lives somewhere that is harder to access.

The schedule lives in the database because scheduling is a coordination problem. It needs uniqueness constraints, it needs to answer “what’s in the 10am slot on Tuesday,” and it needs to do that without me parsing 241 files. A database is genuinely better at that. It just isn’t better at holding prose.

A Database Can Be Useful Without Being Authoritative

I think there’s a reflex where adding a database feels like promoting the data into it. You put the posts in Postgres and now Postgres is where posts are.

It doesn’t have to work that way. A database can be a query layer over data that lives somewhere else, and that’s a completely respectable job. Indexes, joins, counts, “show me every post tagged local-first published before June.” All of that is worth having, and none of it requires the database to be the authority.

The test I use: if I deleted the database right now, what would I lose forever?

For me, it would be the scheduling state because that’s what I put in the database. The important thing is I wouldn’t lose a single word that I’ve written. Every post would still be in a directory. This choice is deliberate.

It’s easy for the database to become a Cache and not an authoritative source.

What Should Happen When They Disagree

If you have documented your authoritative source, then the disagreements stops becoming a crisis, and it just is a routine. Resolution event

You should be able to rebuild it. There should be nothing to decide. The decision is documented and how you resolve conflicts. Just depends on. Which authoritative source owns Which s segment of your data?

In my case, there’s actually a third authoritative source, and that’s the remote blog system that hands back an ID every time I schedule a new post.

So, this is totally fine if you pick the authority at the field level and Document that decision to prevent trip-ups in the future.

Your files and your database shouldn’t be arguing, all it requires is a bit of planning.

I’d appreciate a follow. You can subscribe with your email below. The emails go out once a week, or you can find me on Mastodon at @[email protected].

Databases Software-development Architecture Local-first Data-modeling