- TSSD open binary data format SPEC
author: Brook Zhou(TSSDorg#hotmail.com)
Version: 1
Date: 2026-06-08T07:20:50.52Z UTC
TSSD is an open binary data format, designed for object/struct/class serialize, transfer, exchange or storage.
TSSD is short name for [Type][Size][Size][Data].
TSSD designed simple, cross-platform and high performance
TSSD data are Little Endian
TSSD programing library should implement both of Big and Little Endians.
More backgroud of design in FAQ
L3:|---Magic---|---minor|---major---|---FID---|---Types---|---TID---|---Info---|--payload--|--checksum--|
L2:|-----------head-----------------|-------------schema-----------------------|--payload--|--checksum--|
L1:|-------------------------------Heads---------------------------------------|--Payload--|--Checksum--|
L0:|-----------------------------------------------Fragment---------------------------------------------|
TSSD data are organized by the fragments,
marshal a big object into one big fragment for storage, while mulit fragments with decend size for transfer.
we can describe a fragment in 4 levels, from bottom to up, with more and more details as above picture.
1.1 L1: Heads and Payload need be checksumed, the result will append as Checksum, validate first when the receiver get the fragment.
the couple(sender and receiver) should share the checksum algorithem, which depends on the user and library implement.
1.2 L2: TSSD Heads contain a head and a TSSD schema
1.3 L3: head contain a fixed(5bytes) Magic string: "TSSDV", with 2 bytes TSSD format version, current minor is 1, major is 0, will upgrad in future as this SPEC update.
1.4 L3: TSSD schema contains FID(int16), Types(string), TID(string) and Info(string).
-
Types is short for TSSD Type Schema, is the skeleton of the Payload, it marks which class/struct data, work as a class ID.
-
TID is the ID of current payload, it marks which object data, work as an object ID.
-
FID is the frament ID, range [1, 2, 3,...(N-1), -N], begin with 1, and negative -N marks an ending fragment.
-
Info is reserved for user addtional data.
if a small object marshaled within a fragment, FID should be -1, the first one also the last one.
if a large object marshaled within multi fragments(share the TID), receiver collects all fragments by the Types, TID and FID,
Types is the key of the payload, unmarshal the payload only when your local struct schema Types match with it.
will describe the TSSD data format and Types one by one below.
paylaod is the result of user marshaling an object, an continuity sequential bytes data.
normally TSSD data leading with a TSSD Ttype and following real data.
TSSD expands all fields of an object DFS(deep first search) and collect all the Ttype as the raw Types, as a skeleton, all data of a struct schema should share the same Types.
when marshal finish, user save the payload into one or more fragments, and fill the Types to the Fragment.Schema.Types, so as the TID, FID. but normally we don't need save the full content of Types, a digest or substing of digest(such as MD5) is enough to compare or match.
There are 2 style TSSD formats, fixed-length data and dynamic length data.
a). fixed length data, format [Ttype][data], describe in table 1.
b). dynamic length data:[Ttype][SizeT][SizeA][Data] (our spec name TSSD name for it).
-
SizeT(4bytes signed number): Total size of current chunk, including the following SizeA and Data, but excluding itself.
-
SizeA(2bytes signed number): Additional size, depends the Ttype, describe in table 2.
Note: All SizeT or SizeA are both signed number in little endian
sizet same with SizeT, sizea same with SizeA, both of them print on little endian like that:
sizet 4bytes => [123][456][0][0], but this doc print sizet=[123][456] for short.
so as the sizea 2bytes => [123][0] print as [sizea=123]
Note: Ttype is 1 int8(char) only, default is positive(>0), negative(<0) means omit data, TSSD header always positive
E.g. : [Tint32] means a 32bits number, [-Tint32] means the type is Tint32, but this field unset, just as null in Database.
fixed length data are primary number, the Ttype hints the length
Format: [Ttype][data], data is in Little Endian
Types: [Ttype]
Table 1:
| Ttype | value | CPP type | length(bytes) | range(hex) |
|---|---|---|---|---|
| Tbool | 11 | bool | 1 | true/false |
| Tint8 | 12 | char | 1 | [-0x80,0x7F] |
| Tuint8 | 13 | unsigned char | 1 | [0, 0xFF] |
| Tint16 | 14 | short | 2 | [−0x8000,0x7FFF] |
| Tuint16 | 15 | unsigned short | 2 | [0,0xFFFF] |
| Tint32 | 16 | int32_t | 4 | [-0x80000000,0x7FFFFFFF] |
| Tuint32 | 17 | uint32_t | 4 | [0,0xFFFFFFFFFFFFFFFF] |
| Tint64 | 18 | int64_t | 8 | [-0x8000000000000000,0x7FFFFFFFFFFFFFFF] |
| Tuint64 | 19 | uint64_t | 8 | [0,0xFFFFFFFFFFFFFFFF] |
| Tfloat32 | 20 | float | 4 | 32-bit IEEE 754 |
| Tfloat64 | 21 | double | 8 | 64-bit IEEE 754 |
| fixed length data is simple, just a Ttype following a Little Endian data, Ttype will describe the length data. |
Format: [Ttype][sizet][sizea][data]
Types of the dynamic length data depends the Ttype, will descirbe later one by one.
(Table 2):
| Ttype | value | following data | sizea desc |
|---|---|---|---|
| Tstring | 22 | [sizet][chars] | |
| Ttime | 23 | [Tstring][sizet][RFC3339Nano] | |
| Tenum | 24 | [Tstring][sizet][enum-str] | |
| Tarray | 25 | [sizet][sizea][data] | |
| Tarraym | 26 | [Ttype][sizet][sizea][data] | count of array elements |
| Tobject | 27 | [sizet][sizea][data] | count of the object(struct) fields |
| Tdict | 28 | [sizet][sizea][data] | count of the dict(map) nodes |
| Tdictk | 29 | [Ttype][...] | |
| Tdictv | 30 | [Ttype][...] | |
| Traw | 31 | [sizet][data] | |
| Tschema | 83('M') | [Tobject][sizet][sizea][data] | Tschema is a struct within 4 fields |
| Thead | 84('T') | ["SSD"] | 4 bytes MagicHead: "TSSD" |
| Tversion | 86('V') | [minor][major] | now [1][0] |
| Tuser | 0x7F | [sizet][user-define-data] |
TSSD dynamic length data format and Types detaill describe below.
basic dynamic length data format including Tstring, Ttime, Tschema, Traw, Tuser:
format: [Tstring][sizet][data]
Types: [Tstring]
Tstring is for text string, follow a string length, then the string data
string("Hello TSSD") => [Tstring][sizet=10]{"Hello TSSD"}
format: [Ttime][Tstring][sizet][data]
Types: [Ttime][Tstring]
Ttime presents timestamp, suggest within RFC3339Nano format string
[Ttime][Tstring][sizet=39]{"2023-06-08 11:34:50.371381984 +0000 UTC"}
format: [Tenum][Tstring][sizet][data]
Types: [Tenum][Tstring]
Tenum presents for enum value, suggest within string
it prensents for enum string
[Tenum][Tstring][sizet=9]{"Color.Red"}
format: [Traw][sizet][data]
Types: [Traw]
Traw means raw data, TSSD just forward it without parse.
format: [Tuser][sizet][data]
Types: [Tusr]
Tuser means user define data, TSSD process as Traw now.
Tarray, Tarraym, Tobject, Tdict are composed struct objects, that means they are composed from some of fixed or dynamic length data. all of them can embed within others or themself. they will expand one by one when marshaling. the basic format is as that: [Ttype][SizeT][SizeA][data] list their format detail below with a sample and its explain
format: [Tarray][sizet][sizea][data]
Types: [Tarray][Ttype of the data]
Tarray presents static or dynamic array, sizea means element count in the array.
string[] = {"f", "bar"}
=>
data: [Tarray][sizet=12][sizea=2][Tstring][sizet=1]{"f"}[Tstring][sizet=3]{"bar"}
Types: [Tarray][Tstring]
| value | length(bytes) | desc |
|---|---|---|
| [Tarray] | 1 | array begin |
| [sizet=12] | 2 | array total size: 12 |
| [sizea=2] | 2 | 2 elements |
| [Tstring] | 1 | first element is string |
| [sizet=1] | 2 | the string contains 1 byte |
| {"f"} | 1 | string[0] content: "f" |
| [Tstring] | 1 | second string begin |
| [sizet=3] | 2 | the string contains 3 byte |
| {"bar"} | 3 | string[1] content: "bar" |
format: [Tarraym][Ttype][sizet][sizea][data]
Types: [Tarraym][Ttype]
Tarraym means repeat of the the following Ttype, sizet means total bytes
sizea means element count.
int16[]= {123, 456, 789}
=>
[Tarray][sizet=11][sizea=3][Tint16][123][Tint16][456][Tint16][789]
can also marshaled in Tarraym:
=>
[Tarraym][Tint16][sizet=8][sizea=3][123][456][789]
Tarraym is for POD date only, compare with Tarray, it omit the repeat element Ttype to save storage. element with dynamic length field should marshal in Tarray.
format: [Tobject][sizet][sizea][data]
Types: [Tobject][sizea][Ttype field1][...]
Tobject means a object/struct/class, sizea means fields count in struct, data need expand fields DFS one by one.
object's Types include the field count number after Tobject
struct { int32(123), string("foobar") }
=>
[Tobject][sizet=15][sizea=2][Tint32][123][Tstring][sizet=5]{"foobar"}
Types: [Tobject][2][0][Tint32][Tstring]
| value | length(bytes) | desc |
|---|---|---|
| [Tobject] | 1 | object begin |
| [sizet=15] | 2 | object total size: 15 |
| [sizea=2] | 2 | 2 fields |
| [Tint32] | 1 | first field is int32 |
| [123] | 4 | int32 value |
| [Tstring] | 1 | second field string |
| [sizet=5] | 2 | the string contains 5 byte |
| {"foobar"} | 5 | string content: "foobar" |
format: [Tdict][sizet][sizea]{[Tdictk][k1][Tdictv][v1]}{[Tdictk][k2][Tdictv][v2]}...
Types: [Tdict][Tdictk][Types of key][Tdictv][Types of the value]
Tdict expand and mashaled as key value pair array
Tdictk marks a key begin and Tdictv marks a value begin
the real data of k1, v1, k2, v2 begin with another Ttype and follow above rules
sizea means map node count.
Note: Tdict no guarantee the key order, not conflict or unique. but the library should do.
map{
{int64(234): string("hello")},
{int64(678): string("world!")},
} =>
[Tdict][sizet=41][sizea=2][Tdictk][Tint64][234][Tdickv][Tstring][sizet=5]{"hello"}[Tdictk][Tint64][678][Tdictv][Tstring][sizet=6]{"world!"}
| value | length(bytes) | desc |
|---|---|---|
| [Tdict] | 1 | dict(map) begin |
| [sizet=41] | 2 | dict total size: 41 |
| [sizea=2] | 2 | 2 map nodes |
| [Tdictk] | 1 | a node key begin |
| [Tint64] | 1 | first map node key type |
| [234] | 8 | int64 value |
| [Tdictv] | 1 | a node value begin |
| [Tstring] | 1 | first node value type |
| [sizet=5] | 2 | string length: 5 |
| {"hello"} | 5 | string content: "hello" |
| [Tdictk] | 1 | a node key begin |
| [Tint64] | 1 | second map node key type |
| [678] | 8 | int64 value |
| [Tdictv] | 1 | a node value begin |
| [Tstring] | 1 | second node value type |
| [sizet=6] | 2 | string length: 6 |
| {"world!"} | 6 | string content: "world!" |
Tdictk and Tdictv are used for mark map's key and value begin, the real data format determinted by the following Ttype after them see example at 4.4 Tdict
Tschema is a struct for TSSD meta data, define as that:
struct Schema {
std::int16_t FID; // Fragment ID: [1,2,3...N-1, -N], begin with 1, and -N means an ending fragment
std::string Types; // TSSD Ttype Schema
std::string TID; // TID is the ID of TSSD payload data
std::string Info; // reserve for user additional data info
};format: [Tschema][Tobject][sizet][sizea=4][data]
Schema is meta info for payload within fragment(describe 5).
Fragment ID is from 1 to -N, means N-th fragment, and negative means an ending fragment.
TID is ID of this payload data.
Types is short for TSSD Ttype Schema, skeleton of the payload, work as an class ID.
user can fill the full of the data Types, or digest of the Types(such as MD5), or even substring of the digest(such as 10 chars of leading MD5).
TSSD couple(sender and receiver) should share the Types algorithm, TSSD receiver validate if the Types match with local schema Types, block unmarshal if not.
Tschema(struct {-1, "a1b2c3", "123", "" })
=> [Tschema][Tobject][sizet=23][sizea=4][Tint16][-1 in 2 bytes][Tstring][6]["a1b2c3"][Tstring][3]["123"][Tstring][0]
All TSSD data are wrapped into fragments, frament Format: [header][data][checksum] fragment struct describe as below
struct Fragment {
Header head; // TSSD MAGIC AND Version
Schema schema; // define 4.6
std::byte Payload[]; // TSSD paylod data
std::byte Checksum[]; // disgest of all the Fragment bytes
};receiver(reader) collects all Fragments by the Schema.Types and Schema.TID, and order by the Schema.FID, will make up a whole TSSD Data, then can be unmarshaled to a user object.
Format: [Thead='T'][header="SSD"][Tversion='V'][minor=1][major=0][Tschema='M'][Tobject][sizet][sizea][schema-data]
| item | value | length | desc |
|---|---|---|---|
| Thead | ['T'] | 1byte | TSSD header begin |
| header | ['S','S','D'] | 3bytes | 3 fixed charactar |
| Tversion | ['V'] | 1byte | version begin |
| TSSD version | [minor][major] | 2bytes | TSSD minor major version, current: [1][0] |
| Tschema | ['M'] | 1byte | schema begin |
| Tobject | [Tobject] | 1byte | object(struct) begin |
| sizet | 4bytes | schema marshal total bytes | |
| sizea | [4][0] | 2bytes | schema object fields(4) |
| schema data | [xxxx] | sizet | schema object data in TSSD format |
"TSSDV" is the magic header in short and marks a TSSD fragment begin
TSSD version: minor version 1, major version 0. may update in future.
payload work as Tarraym with a TSSD header: [Tarraym][Tuint8][sizet][sizea][payload content]
payload content is the real use data within TSSD format, nomally we call payload the user data only, without the Tarraym header.
current fragment checksum data, is also marshal into Tarraym: [Tarraym][Tuint8][sizet][sizea][checksum data]
Note: checksum should include fragment header and payload data.
if checksum data is empty, skip the checksum validation.
- TSSD library marshal user's object, and fill the Schema info, and put in a fragment, to store or tranfer on the network.
- TSSD receiver(library) unmarshal the raw byte into Fragments and validate the checksum and validate if the Types match with local.
- receiver collects all the fragments by the Schema's Types, TID, FragmentID.
- unmarshal the payload object.
expand an object into TSSD folloing DFS(deep first search), so as the Types produce algorithem. an example for Types:
struct StructA {
bool vbool;
std::string strs[2];
};
struct StructB {
std::int32_t vint32;
std::map<std::int16_t, StructA> mp;
StructA va;
};the full raw Types of StructB should be:
[Tobject][3][0][Tint32][Tdict][Tdictk][Tint16][Tdictv][Tobject][2][0][Tbool][Tarray][Tstring][Tobject][2][0][Tbool][Tarray][Tstring]
simple description as following table:
| Bytes | Descript |
|---|---|
| [Tobject][3][0] | an StructB object with 3 fields |
| [Tint32] | StructB.vint32 |
| [Tdict][Tdictk][Tint16] | an map with key type int16 |
| [Tdictv][Tobject][2][0][Tbool][Tarray][Tstring] | map's value(StructA) |
| [Tobject][2][0][Tbool][Tarray][Tstring] | StructB.va |
- be careful with the type that language or platform dependable, if you need TSSD data share cross platform.
- int: length may vary in 32bits and 64bit.
- unsigned number: Java does't implement it.
- Tdict(map) key attribute may vary with the language: Go accepts basic type only, while CPP can accept complex struct.
- TSSD can unmarshal within diff local type sometimes, such as sender marshal a std::list into Tarray, and receiver unmarshal into std::vector