|
JSON Voorhees
Killer JSON for C++
|
JSON Voorhees is a JSON library written for the C++ programmer who wants to be productive in this modern world. This one targets C++23 for developer-friendliness, a reasonably fast parser, and no dependencies beyond a compliant compiler and standard library. It is hosted on GitHub and sports an Apache License, so use it anywhere you need.
Features include (but are not necessarily limited to):
value should not feel terribly different from a C++ Standard Library containeroperator<<parsevalue is 16 bytes on a 64-bit platform)value, using deserialize<T>serialize, or into a value using to_json[[nodiscard]], so dropping the answer to a question you asked is a warningJSON Voorhees is designed with ease-of-use in mind. So let's look at some code!
The central class of JSON Voorhees is the jsonv::value, which represents a JSON AST. Putting values of different types is easy.
Output:
If that isn't convenient enough for you, there is a user-defined literal _json in the jsonv namespace you can use:
JSON is dynamic, which makes value access a bit more of a hassle, but JSON Voorhees aims to make it not too horrifying for you. A jsonv::value has a number of accessor methods named things like as_integer and as_string which let you access the value as if it was that type. But what if it isn't that type? In that case, the function will throw a jsonv::kind_error with a bit more information as to what rule you violated.
Output:
You can also deal with container types in a similar manner that you would deal with the equivalent STL container type, with some minor caveats. Because the value_type of a JSON object and JSON array are different, they have different iterator types in JSON Voorhees. They are named object_iterator and array_iterator. The access methods for these iterators are begin_object / end_object and begin_array / end_array, respectively. The object interface behaves exactly like you would expect a std::map<std::string,jsonv::value> to, while the array interface behaves just like a std::deque<jsonv::value> would.
Output:
The iterator types work. This means you are free to use all of the C++ things just like you would a regular container. To use a ranged-based for, simply call as_array or as_object. Everything from <algorithm> and <iterator> or any other library works great with JSON Voorhees.
Output:
Usually, the reason people are using JSON is as a data exchange format, either for communicating with other services or storing things in a file or a database. To do this, you need to encode your json::value into an std::string and parse it back. JSON Voorhees makes this easy for you.
Output:
If you are paying close attention, you might have noticed that the value for the "infinity" looks a little bit more null than infinity. This is because, much like mathematicians before Anaximander, JSON has no concept of infinity, so it is actually illegal to serialize a token like infinity anywhere.
By default, when an encoder encounters an unrepresentable value in the JSON it is trying to encode, it outputs null instead. If you wish to change this behavior, implement your own jsonv::encoder (or derive from jsonv::ostream_encoder).
If you ran the example program, you might have noticed that the return code was 1, meaning the value you put into the file and what you got from it were not equal. This is because all the type and value information is still kept around in the in-memory obj. It is only upon encoding that information is lost.
Getting tired of all this compact rendering of your JSON strings? Want a little more whitespace in your life? Then jsonv::ostream_pretty_encoder is the class for you! Unlike our standard compact encoder, this guy will put newlines and indentation in your JSON so you can present it in a way more readable format.
Compile that code and you now have your own little JSON prettification program!
Not everything you want to write starts out as a jsonv::value. A jsonv::writer writes a document one token at a time into any encoder, the pretty one included, and jsonv::serialize writes a C++ object as one value wherever the writer is:
Output:
The writer writes the commas and the colons, and refuses a token which does not belong where it is – a key outside an object, say – before the encoder ever sees it. Each call to serialize writes one more element into the array the writer has open, so the same loop could write a million names without ever holding them all in a jsonv::value. serialize works for any type a jsonv::formats knows – here, jsonv::formats::global() – and teaching one about your own types is what the next section is about.
Most of the time, you do not want to deal with jsonv::value instances directly. Instead, most people prefer to convert JSON into their own strong C++ class or struct. JSON Voorhees provides utilities to make this easy for you to use. At the end of the day, you should be able to create an arbitrary C++ type with jsonv::deserialize<my_type>(text) and turn one back into JSON text with jsonv::serialize(my_instance) – or into a jsonv::value with jsonv::to_json(my_instance).
Let's start with converting JSON into C++ types with jsonv::deserialize<T>.
Output:
The first three deserialize from JSON text, which is the spelling to reach for when text is what you have: the C++ value is read straight out of the text, and no jsonv::value is built along the way. That does mean a C++ string handed to deserialize is JSON text rather than a JSON string, so deserialize<std::string>(R"("Hello!")") is Hello! while deserialize<std::string>("Hello!") fails to parse. The last deserializes from a jsonv::value, which is the spelling for JSON you have already parsed or built. Either way, JSON which does not hold what you asked for – deserialize<int>(R"("one")") – throws a jsonv::deserialization_error saying what was found instead.
Overall, this is not very complicated. We did not do anything that could not have been done through a little use of parse and the as_ accessors like as_integer. So what is this deserialize giving us?
The real power comes in when we start talking about jsonv::formats. These objects provide a set of rules to encode and decode arbitrary types. So let's make a C++ class for our JSON object and write a special constructor for it.
Output:
There is a lot going on in that example, so let's take it one step at a time. First, we are creating a my_type object to store our values, which is nice. Then, we gave it a funny-looking constructor:
This is a deserializing constructor. All that means is that it has those two arguments: a jsonv::reader and a jsonv::deserialization_context. The reader is a forward cursor over the JSON, and when the constructor is called it is sitting on the first token of the value to deserialize from – for my_type, the { of an object. From there, the constructor walks the object one key at a time, in whatever order the document wrote them:
Each value it wants is deserialized by deserialize_member, which leaves the reader on the next key, or on the }. A key it does not recognize has its value skipped with next_value, which steps over the whole value in one go, however large it is. Once the closing } has been stepped off as well, the reader is left one position past the object. Every deserializer promises that, because it is where whatever is deserializing around this object carries on from. A key the document leaves out leaves its member as it was initialized.
The jsonv::deserialization_context is what does the work. context.deserialize<T>(from) deserializes a T from the value under the cursor, using the jsonv::formats the deserialization was started with, and leaves the cursor one past that value. When it cannot, it does not throw: it records the problem on the context and returns a std::unexpected. A constructor can only fail by throwing, so deserialize_member throws a jsonv::deserialization_error carrying what the context recorded – taking it with take_problems_since rather than copying it, so the problem is reported once. The path_scope names the member for as long as it is being deserialized, which is what puts a problem with "a" at .a – or at [3].a when the my_type is the fourth element of an array.
Giving up at the first problem is what deserialization does by default. With jsonv::deserialize_options::on_error::collect_all, it carries on past a problem so that it can report as many as it finds, and a deserializer which walks the reader itself has more to do for that to work – see jsonv::deserialization_context::recover. The DSL described below does all of that for you.
A jsonv::deserializer is a type that knows how to read JSON and create some C++ type out of it. In this case, we are creating a jsonv::deserializer_construction, which is a subtype that knows how to call the constructor of a type. There are all sorts of jsonv::deserializer implementations in jsonv/serialization/, so you should be able to find one that fits your needs.
Now things are starting to get interesting. The jsonv::formats object is a collection of jsonv::deserializers, so we create one of our own and add the jsonv::deserializer* from the static function of my_type. The local_formats only knows how to deserialize instances of my_type – it does not know even the most basic things like how to deserialize an int. We use jsonv::formats::compose to create a new instance of jsonv::formats that combines the qualities of local_formats (which knows how to deal with my_type) and the jsonv::formats::defaults (which knows how to deal with things like int and std::string). The formats instance now has the power to do everything we need!
This is not terribly different from the example before, but now we are explicitly passing a jsonv::formats object to the function. If we had not provided format as an argument here, the function would have thrown a jsonv::deserialization_error complaining about how it did not know how to deserialize a my_type.
When the JSON came from a file, an error is more use if it says which file. Build the jsonv::deserialization_context yourself, with the name of the source last, and hand it to deserialize in place of the format:
The "b" is not an int, so this throws a jsonv::deserialization_error reading Deserialization error at my_type.json#.b: Read node of type string when expecting integer. Write out every argument before the name: the nullptr is the user data, and a string in its place would be taken for user data rather than for a name. A context is meant for one document, so make a new one for each file.
If you are coming from JSON Voorhees 1.x, you may be looking for extract_sub. A deserializing constructor used to be handed a whole jsonv::value and pull each member out of it by name, which is what extraction_context::extract_sub did. It is gone because that value is gone: deserialization reads the JSON as it goes rather than building a value first, and a forward cursor has no way to look a key up. Walking the keys, as my_type does, is what replaces it. When random access is genuinely wanted – what one member means depends on another written after it, say – read the object into a jsonv::value and look things up in that, with value::find or value::at_path. A deserializing constructor can still take a const jsonv::value& in place of the reader for exactly this, and is handed the object as a value – read off the reader with jsonv::read_value, if it was not one already. That keeps the extract_sub calls of a 1.x constructor easy to port:
The path_scope is what says where a problem is, since a value has no idea where in the document it came from. That includes a member which is not there at all: looking it up with find rather than value::at is what reports a missing "b" at .b, while the scope is still alive to say so. The std::out_of_range from at would only be caught once the scope had gone, so it would be reported at the object around the member, and it does not say which member it was. context.deserialize<T> throws rather than returning when it is handed a value, so nothing needs handing over. Building that value for every object deserialized is the cost the reader-based constructor avoids.
JSON Voorhees also converts from your C++ structures into JSON, using jsonv::serialize for JSON text and jsonv::to_json for a jsonv::value. It should feel like a mirror of jsonv::deserialize, with similar argument types and many shared concepts. Just like deserialization, both use the jsonv::formats class, but they look up a jsonv::serializer in it to convert from C++ into JSON. Where a deserializer reads from a jsonv::reader, a serializer writes into a jsonv::writer.
Output:
The serializer is handed a jsonv::writer positioned where the my_type goes – at the root here, but just as well as the element of an array or the value of another object's member – and writes exactly one value there. It opens an object, and for each member writes the key and then hands the member to context.serialize, which looks up the serializer for the member's type in the same jsonv::formats and has it write the member where the writer now is. The writer writes the punctuation, and refuses a token which does not belong where it is: a key outside an object, say, or a value inside one with no key before it.
jsonv::serialize writes x as compact JSON text with nothing built in between: no jsonv::value is made along the way, and the members come out in the order the serializer wrote them. jsonv::to_json runs the same serializer into a jsonv::value_encoder and hands back the tree it built, for when a jsonv::value is what you want – to look things up in, as here, or to change before writing it out. A jsonv::value keeps an object's members sorted by key, so the text of one can list them in a different order from serialize; for my_type, they happen to agree. To write the text to a stream rather than into a std::string, there is jsonv::serialize(x, std::cout, format); and jsonv::serialize(x, to, format) writes x into a jsonv::writer of your own, as in Encoding and decoding.
A serializer written for JSON Voorhees 1.x returned a jsonv::value rather than writing one. That still works: jsonv::make_serializer also takes a function of (const jsonv::serialization_context&, const T&), or of just the const T&, which returns a jsonv::value to be written whole. Building that value for every object serialized is the cost the writer avoids.
Does all this seem a little bit manual to you? Creating a deserializer and serializer for every single type can get a little bit tedious. Unfortunately, until C++ has a standard way to do reflection, we must specify the conversions manually. However, there is an easier way! That way is the Serialization Builder DSL.
Let's start with a couple of simple structures:
Let's make a formats for them using the DSL:
What is going on there? The giant chain of function calls is building up a collection of type adapters into a formats for you. The indentation shows the intent – the .member("a", &foo::a) is attached to the type adapter for foo (if you tried to specify &bar::y in that same place, it would fail to compile). Each function call returns a reference back to the builder so you can chain as many of these together as you want to. The jsonv::formats_builder is a proper object, so if you wish to spread out building your type adapters into multiple functions, you can do that by passing around an instance.
The two most-used functions are type and member. type defines a jsonv::adapter for the C++ class provided at the template parameter. All of the calls before the second type call modify the adapter for foo. There, we attach members with the member function. This tells the formats how to encode and deserialize each of the specified members to and from a JSON object using the provided string as the key. The extra function calls like default_value, since and until are just a couple of the many functions available to modify how the members of the type get transformed.
The chain ends with compose_checked, which checks that every type the members refer to – int and std::string here – can be deserialized and serialized once the adapters the DSL built are combined with jsonv::formats::defaults, and then composes the two, just as we composed local_formats by hand earlier.
The formats we built would be perfectly capable of serializing to and deserializing from this JSON document:
Deserializing a bar reads the document just the way the constructor of my_type did: each object's keys are walked in the order the document wrote them, each is handed to the member which claims it, and a key no member claims has its value stepped over unread. What the DSL adds is everything that constructor left out. A member without a default_value is required, a key the document repeats is settled by jsonv::deserialize_options::on_duplicate_key, and jsonv::deserialize_options::on_error::collect_all carries on past a problem to report the rest.
For a more in-depth reference, see the Serialization Builder DSL page.
JSON Voorhees takes a "batteries included" approach. A few building blocks for powerful operations can be found in the algorithm.hpp header file.
One of the simplest operations you can perform is the map operation. This operation takes in some jsonv::value and returns another. Let's try it.
If everything went right, you should see a number:
That is not the most interesting example of using map, but it is enough to get the general idea of what is going on. This operation is so common that it is a member function of value as jsonv::value::map. Let's make things a bit more interesting and map an array...
Now we're starting to get somewhere!
The map function maps over whatever the contents of the jsonv::value happens to be and returns something for you based on the kind. This simple concept is so ubiquitous that Eugenio Moggi named it a monad. If you're feeling adventurous, try using map with an object or chaining multiple map operations together.
Another common building block is the function jsonv::traverse. This function walks a JSON structure and calls a some user-provided function.
Now we have a tiny little program to decompose JSON into jq style path expressions and their values. For example, if you pipe { "bar": [1, 2, 3], "foo": "hello" } into the program:
All of the really powerful functions can be found in algorithm.hpp. My personal favorite is jsonv::merge. The idea is simple: it merges two (or more) JSON values into one.
Output:
You might have noticed the use of std::move into the merge function. Like most functions in JSON Voorhees, merge takes advantage of move semantics. In this case, the implementation will move the contents of the values instead of copying them around. While it may not matter in this simple case, if you have large JSON structures, the support for movement will save you a ton of memory.