C++ JSON Parsing: nlohmann/json, RapidJSON, and Custom Types
The Problem with Naive JSON Parsing
C++ has no built-in JSON support, and rolling your own parser is an invitation to bugs. The common pitfalls when consuming JSON from REST APIs or config files:
- Type mismatch: the API sends
"age": "30"(string) but your code calls.get<int>() - Missing keys:
j["optional_field"]silently inserts a null into the object - Parse errors: malformed JSON crashes the process if exceptions aren’t caught
- Memory bloat: loading a 50MB JSON file into memory all at once
Two libraries dominate C++ JSON parsing: nlohmann/json (ergonomic, STL-like) and RapidJSON (fast, low memory). Others are worth knowing: simdjson parses at very high speed using SIMD instructions but is read-only, Boost.JSON sits between the two in speed and ergonomics and fits projects already on Boost, and Glaze maps structs to JSON at compile time.
The underlying difficulty is the same in every library: JSON is dynamically typed and C++ is not. Every value you read has to be checked at runtime for presence and type before it can become an int or a std::string, and every library has to decide what happens when that check fails (throw, return an error code, or give you undefined behavior). Most production bugs with JSON in C++ come from assuming a shape the input did not have, so this article spends as much time on failure handling as on the happy path.
Library Comparison
| nlohmann/json | RapidJSON | |
|---|---|---|
| API style | STL-like, intuitive | Verbose, explicit |
| Parse speed | Good | Faster (pool allocator, optional in-situ parsing) |
| Memory | Higher | Lower; SAX mode available |
| Custom types | to_json / from_json | Manual mapping |
| Error handling | Exceptions | HasParseError() check |
| Header-only | Yes | Yes |
| Best for | Most projects | High-throughput, large files |
The two represent opposite design goals. nlohmann/json makes a json value behave like an STL container and converts implicitly to and from C++ types, so code reads almost like Python; every object key and value is a separately heap-allocated node, and compile times grow noticeably because the header is large and heavily templated. RapidJSON exposes allocators and ownership explicitly (you pass an allocator when building values, and strings can be copied or referenced), which is more work and more ways to make a mistake, but gives much tighter control over memory. For typical REST handlers and config files, parsing time is small compared with network and database time, and nlohmann’s readability wins.
nlohmann/json
Install via CMake’s FetchContent or copy the single header:
#include <nlohmann/json.hpp>
using json = nlohmann::json;
Parsing a String or File
#include <nlohmann/json.hpp>
#include <fstream>
#include <stdexcept>
using json = nlohmann::json;
// Parse from string
json parseFromString(const std::string& raw) {
try {
return json::parse(raw);
} catch (const json::parse_error& e) {
throw std::runtime_error(
"JSON parse error at byte " + std::to_string(e.byte) + ": " + e.what()
);
}
}
// Parse from file
json parseFromFile(const std::string& path) {
std::ifstream file(path);
if (!file.is_open()) {
throw std::runtime_error("Cannot open file: " + path);
}
try {
return json::parse(file);
} catch (const json::parse_error& e) {
throw std::runtime_error("Parse error in " + path + ": " + e.what());
}
}
json::parse throws parse_error for malformed input, and its message already includes the position, for example [json.exception.parse_error.101] parse error at line 3, column 15: syntax error while parsing object - unexpected '}'; expected string literal. The byte member gives the offset for logging. If exceptions are unwelcome (a hot path, or a codebase compiled without them), json::parse(raw, nullptr, false) returns a value for which is_discarded() is true instead of throwing, and json::accept(raw) only validates. Parsing from an std::ifstream avoids reading the file into a string first; opening it in binary mode is not required, but a UTF-8 BOM at the start of the file is rejected unless the library version skips it (recent 3.x versions do).
Two inputs that surprise people: an empty string or empty file throws parse_error.101 ... unexpected end of input, which is often how a failed HTTP request without a body shows up, and a file with comments (// ...) is invalid JSON unless you pass ignore_comments = true as the fourth argument of parse.
Safe Field Access
The most common mistake is using j["key"] for optional fields — it silently inserts null. Use these patterns instead:
void processApiResponse(const json& j) {
// Required field — throws json::out_of_range if missing
std::string id = j.at("id").get<std::string>();
// Optional field with default
std::string name = j.value("name", "Unknown");
// Check before accessing
if (j.contains("email")) {
std::string email = j["email"].get<std::string>();
}
// Type checking before conversion
if (j.contains("age") && j["age"].is_number_integer()) {
int age = j["age"].get<int>();
}
// Nested object — check each level
if (j.contains("address") && j["address"].is_object()) {
const auto& addr = j["address"];
std::string city = addr.value("city", "");
}
// Array iteration
if (j.contains("tags") && j["tags"].is_array()) {
for (const auto& tag : j["tags"]) {
if (tag.is_string()) {
std::cout << tag.get<std::string>() << '\n';
}
}
}
}
The function takes const json& on purpose. On a non-const json, j["missing"] inserts a null value and returns a reference to it, so a read silently changes the document, and a later dump() contains "missing": null. On a const json, the same expression with a missing key is undefined behavior; the library asserts in debug builds, and a release build may crash. The contains checks above are what make the j["..."] reads safe. find() avoids the double lookup: if (auto it = j.find("email"); it != j.end() && it->is_string()) checks presence and type once.
value("name", "Unknown") looks like the complete answer for optional fields, but it only covers the missing-key case. If the key exists with the wrong type ("name": 42 or "name": null), value throws type_error rather than returning the default. APIs that send null for “not set” are common, so for fields that can be null, check is_null() explicitly. Also note that is_number_integer() is false for 30.0, and get<int>() on a number that does not fit in int silently truncates or wraps rather than throwing, so range-check values that come from outside.
Error Taxonomy
try {
auto j = json::parse(raw_input);
// type_error: wrong type assumed
int x = j["count"].get<int>(); // throws if "count" is a string
} catch (const json::parse_error& e) {
// Malformed JSON — log e.byte for the failure location
log_error("Parse failed at byte {}: {}", e.byte, e.what());
} catch (const json::type_error& e) {
// Type mismatch during .get<T>() or operator usage
log_error("Type error: {}", e.what());
} catch (const json::out_of_range& e) {
// .at("key") when key doesn't exist
log_error("Missing field: {}", e.what());
}
All three derive from nlohmann::json::exception, which in turn derives from std::exception, and each carries a numeric id (such as 302 for a type mismatch or 403 for a missing key) that is stable across versions and useful in logs and tests. In this example, j is non-const, so if count is missing, j["count"] inserts null and get<int>() throws type_error.302: type must be number, but is null, not out_of_range; the error message then misleadingly suggests a type problem when the field is simply absent. Using .at("count") gives the accurate out_of_range.403: key 'count' not found.
In a request handler, the boundary is the right place for this try: catch the JSON exceptions once, turn them into a 400 response with the message, and let the rest of the code assume a validated structure. Catching them deep inside business logic tends to swallow errors that should have rejected the request.
Custom Type Serialization
For structs, define to_json and from_json as free functions:
struct User {
std::string id;
std::string email;
int age{0};
std::optional<std::string> phone; // optional field
};
void to_json(json& j, const User& u) {
j = json{
{"id", u.id},
{"email", u.email},
{"age", u.age},
};
if (u.phone.has_value()) {
j["phone"] = u.phone.value();
}
}
void from_json(const json& j, User& u) {
u.id = j.at("id").get<std::string>();
u.email = j.at("email").get<std::string>();
u.age = j.value("age", 0);
if (j.contains("phone") && j["phone"].is_string()) {
u.phone = j["phone"].get<std::string>();
}
}
// Usage — automatic conversion
User user = json::parse(raw).get<User>();
json j = user; // serialize back to JSON
std::cout << j.dump(2) << '\n'; // pretty-print with 2-space indent
to_json and from_json are found by argument-dependent lookup, so they must be declared in the same namespace as User; defining them in a different namespace produces a long compile error ending in something like could not find to_json() method in T's namespace. With them in place, get<User>(), assignment to json, and even std::vector<User> or std::map<std::string, User> conversions work automatically, because the library composes conversions for standard containers.
The mapping choices here are deliberate: id and email are required and use at(), so a missing one throws out_of_range naming the key; age falls back to 0; phone is written only when present, so the output does not contain "phone": null. That asymmetry between required and optional fields is precisely the information the macro shortcut below cannot express.
Macro Shortcut for Simple Structs
struct Config {
std::string host;
int port{8080};
bool debug{false};
};
// Generates to_json and from_json automatically
NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE(Config, host, port, debug)
// Now: Config cfg = json::parse(raw).get<Config>();
The macro generates from_json using at() for every listed member, so every field becomes required: a config file without "debug" throws out_of_range even though the struct has a default of false. Since version 3.11, NLOHMANN_DEFINE_TYPE_NON_INTRUSIVE_WITH_DEFAULT uses the member’s current value when a key is missing, which is usually what config structs want. The _INTRUSIVE variants go inside the class and can access private members. Like hand-written functions, the macro must appear in the namespace of the type.
RapidJSON
RapidJSON is faster and uses less memory, at the cost of a more verbose API:
#include <rapidjson/document.h>
#include <rapidjson/error/en.h>
#include <rapidjson/filereadstream.h>
using namespace rapidjson;
void parseWithRapidJson(const std::string& raw) {
Document doc;
doc.Parse(raw.c_str());
// Always check for parse errors
if (doc.HasParseError()) {
fprintf(stderr, "Parse error at offset %zu: %s\n",
doc.GetErrorOffset(),
GetParseError_En(doc.GetParseError()));
return;
}
// Safe field access
if (doc.HasMember("name") && doc["name"].IsString()) {
std::string name = doc["name"].GetString();
}
if (doc.HasMember("count") && doc["count"].IsInt()) {
int count = doc["count"].GetInt();
}
// Array
if (doc.HasMember("items") && doc["items"].IsArray()) {
const auto& items = doc["items"];
for (SizeType i = 0; i < items.Size(); ++i) {
if (items[i].IsString()) {
printf("%s\n", items[i].GetString());
}
}
}
}
RapidJSON does not throw. A parse failure sets an error code and offset on the document, and access to a missing member or a value of the wrong type is guarded only by RAPIDJSON_ASSERT, which is assert() by default: a debug build aborts, and a release build reads garbage. That is why every access above is preceded by HasMember and an Is...() check. doc.FindMember("name") returns an iterator and avoids looking the key up twice, which matters because member lookup in RapidJSON objects is a linear scan. GetString() returns a pointer into the document’s memory, valid only as long as doc lives; copying into std::string, as the name line does, is the safe default.
IsInt() is stricter than it looks: it is false for values that only fit in int64_t or uint32_t, and false for 30.0. RapidJSON distinguishes Int, Uint, Int64, Uint64 and Double, so check the widest type you are prepared to accept.
File Streaming (Low Memory)
For large files, stream directly rather than loading into a std::string:
#include <rapidjson/filereadstream.h>
#include <cstdio>
void parseLargeFile(const char* path) {
FILE* fp = fopen(path, "rb");
if (!fp) return;
char readBuffer[65536];
FileReadStream is(fp, readBuffer, sizeof(readBuffer));
Document doc;
doc.ParseStream(is);
fclose(fp);
if (doc.HasParseError()) {
fprintf(stderr, "Error: %s\n", GetParseError_En(doc.GetParseError()));
}
}
What streaming saves here is the copy of the raw text: the file is read through a 64 KB buffer instead of being loaded into a std::string first. The resulting Document still holds the entire parsed tree in memory, so peak memory is roughly the size of the DOM, which for RapidJSON is usually of the same order as the file itself. To process a file that does not fit in memory, use the SAX interface (Reader::Parse with a handler whose Key, String, Int, StartObject and similar callbacks receive each value as it is read), where memory stays roughly constant regardless of document size. SAX code is harder to write because you have to track your position in the structure yourself, so it is worth it only for genuinely large inputs or for streaming pipelines.
Performance Comparison
What each library does while parsing a large document explains most of the difference:
| Library | What parsing builds | Memory behavior |
|---|---|---|
| nlohmann/json | A tree of json values; objects are std::map<std::string, json> by default | One heap allocation per key string, container, and node, so peak memory is well above the file size |
| RapidJSON DOM | A tree allocated from a memory pool (MemoryPoolAllocator); in-situ mode can reuse the input buffer for strings | Far fewer allocations, freed all at once |
| RapidJSON SAX | No tree at all; your handler gets callbacks per value | Memory stays roughly constant regardless of document size |
The exact timings depend on the document shape (many small objects vs. long strings), compiler, and flags, so benchmark with your own payloads before switching.
Rule: use nlohmann/json unless profiling shows JSON parsing as a bottleneck. If you process megabyte-scale JSON frequently, switch to RapidJSON SAX.
Production Checklist
- Always catch
parse_errorwhen parsing untrusted input (API responses, config files, user uploads) - Never use
j["key"]for optional fields — it silently inserts null and corrupts your object - Cap input size before parsing — reject payloads over your maximum (e.g., 1MB for API requests)
- Log the byte offset from
parse_error.byteso you can diagnose truncated or malformed messages - Thread safety: concurrent reads of a
const jsonare fine, but any mutation (including a non-constj["key"]read that inserts) needs synchronization — copy or parse per thread - Validate required fields with
.at()so missing required data throws immediately rather than silently defaulting - Encode UTF-8 — both libraries expect UTF-8; validate incoming byte sequences from untrusted sources. In nlohmann/json,
dump()throwstype_error.316: invalid UTF-8 bytewhen a string contains invalid bytes (for example text read from a Latin-1 file);dump(-1, ' ', false, json::error_handler_t::replace)substitutes U+FFFD instead - Limit nesting depth — both parsers recurse on nested arrays and objects, so deeply nested hostile input can exhaust the stack; cap payload size and, where the library allows, nesting depth
Choosing and using a JSON library
- nlohmann/json is the right default — readable, STL-compatible, and supports automatic custom type serialization via
to_json/from_json - RapidJSON wins on raw performance and memory; use its SAX API for multi-MB files
- Use
value("key", default)orcontains()for optional fields — never nakedj["key"] - The error hierarchy matters:
parse_error(bad JSON),type_error(wrong type),out_of_range(missing required key) - Define
to_json/from_jsonor useNLOHMANN_DEFINE_TYPE_NON_INTRUSIVEto keep struct mapping maintainable
Frequently Asked Questions (FAQ)
Q. Why do I get type_error even though the field exists?
A. get<T>() and implicit conversions throw nlohmann::json::type_error when the stored JSON type does not match the requested C++ type, for example a number sent as the string “42” or a null where you expect a string. value("key", default) throws in that case too, because it substitutes the default only when the key is missing, not when the type is wrong. For untrusted input, check is_number() / is_string() before converting, or catch type_error at the boundary and turn it into a validation error.