The schema: field numbers, not field names

Encoding each field by a stable NUMBER (not its name) and tagging its wire type lets a reader skip fields it doesn't recognize — which is exactly what lets old and new code talk while the schema evolves.

Previously

We can pull one whole message off the wire, but it's a meaningless blob until both sides agree how to read it — so we hand the bytes a shared schema instead of hand-rolled JSON.

Scene 03

The schema: field numbers, not field names

  1. Watch
  2. Try it
  3. Predict
  4. Capture
.proto · schema V1message Greeting { string name = 1;}SCHEMA MATRIXwriter: V1reader: V1fields keyed by numberENCODE → BYTES ON THE WIREtag (field #) · length · valuename = 1#1tag3lenAdastringREADER (V1) DECODES BY NUMBERnameAdaevery field recognized — clean decodetag-length-value: each field is laid out as [tag · field number] [length] [value] · the tag is a varint (field number << 3 | wi…
this is the schema — the .proto contract for the message
What to watch for

Both sides of an RPC need to agree on what the bytes inside a frame MEAN — and in a fleet of services written in many languages, that agreement can't be a comment in a README. We write it down once. A schema is a typed contract for a message: a short file (a .proto) that lists each field, its type, and a stable number. A code generator turns that one file into typed client and server code in every language — no hand-rolled JSON, no per-team drift. Watch Greeting{name="Ada"} encode to bytes: the field becomes a tag block, then a length, then the UTF-8 'Ada'. Notice what the tag block says — it names the field number (#1), not the word 'name'. That single choice is the whole scene.

Continue unlocks when the animation finishes.
Implementation

Highlighted lines are the ones running in the diagram right now.

Codec.tag
the tag packs the field number with its wire type
def tag(field_number, wire_type):
# one varint: number in the high bits, type in low 3
return varint(field_number << 3 | wire_type)
WIRE_VARINT = 0 # int32, bool, enum
WIRE_LEN = 2 # string, bytes, sub-message
Writer.encode
each field becomes tag, then length, then value
def encode(msg):
out = bytes()
out += tag(1, WIRE_LEN) # field #1 = name
out += varint(len(msg.name))
out += utf8(msg.name)
if schema >= v2: # v2 adds field #2
out += tag(2, WIRE_VARINT) # field #2 = priority
out += varint(msg.priority)
return out
Reader.decode
decode by number; an unknown number is stepped over
def decode(buf):
while buf:
field_number, wire_type = untag(buf.read_varint())
if field_number in self.schema:
msg[field_number] = read_value(buf, wire_type)
else: # number we never defined
length = buf.read_varint()
buf.skip(length) # step past it, no crash
return msg

Where this sits in Build a gRPC-style RPC framework

Scene 03 of 14, in the The wire act — Bytes, frames, and the schema that gives them meaning.. A schema keys each field by a stable number, not its name — so an old reader can skip a field it doesn't know and still decode the rest. Never reuse a field number.

Up next. Our typed Greeting{name="Ada"} bytes are ready to send — but on one raw TCP connection we're back to one-call-at-a-time; next we need a transport that runs many calls down one pipe at once.

All 14 scenes in Build a gRPC-style RPC framework · Every curriculum

Built with Arqly
Every scene in Build a gRPC-style RPC framework builds on the one before it.All 14 Build a gRPC-style RPC framework scenes