JSON vs YAML vs TOML vs XML vs INI: Picking a Config Format and Converting Between Them
Key takeaways
JSON, YAML, TOML, XML and INI side by side: syntax, how each is parsed in Python and JavaScript, and the traps that actually break configs, such as YAML's implicit typing, INI files without sections, JSON's missing comments and unsafe YAML loading, plus tools for converting and validating.
Introduction: The Importance of Configuration File Formats
Every programming project uses configuration files. package.json, docker-compose.yml, nginx.conf, README.md - each format has different characteristics and purposes. Choosing the right format greatly improves code readability and maintainability.
Formats covered in this article:
- JSON (JavaScript Object Notation)
- YAML (YAML Ain’t Markup Language)
- XML (eXtensible Markup Language)
- TOML (Tom’s Obvious, Minimal Language)
- INI (Initialization File)
- Markdown
- Others (Properties, HCL, Jsonnet)
Often the format is not yours to choose: npm reads package.json, Cargo reads Cargo.toml, Kubernetes takes YAML. The choice is real when you design your own application’s config, and there the questions that matter are whether humans edit the file (comments, readability), whether values must keep exact types (a version string must not become a float), and how easy it is to break the file with a small typo. The sections below go format by format, with the traps that cause real outages, not just syntax.
JSON (JavaScript Object Notation)
What is JSON?
JSON is a lightweight data interchange format based on JavaScript object notation. It’s easy for humans to read and easy for machines to parse.
JSON Basic Syntax
{
"name": "John Doe",
"age": 30,
"email": "[email protected]",
"isActive": true,
"balance": 1234.56,
"tags": ["developer", "designer"],
"address": {
"street": "123 Main St",
"city": "Seoul",
"zipCode": "12345"
},
"projects": [
{
"id": 1,
"name": "Project A",
"status": "active"
},
{
"id": 2,
"name": "Project B",
"status": "completed"
}
],
"metadata": null
}
JSON Data Types
{
"string": "Hello, World!",
"number": 42,
"float": 3.14159,
"boolean": true,
"null": null,
"array": [1, 2, 3, "mixed", true],
"object": {
"nested": "value"
}
}
JSON Pros and Cons
flowchart TB
JSON[JSON]
subgraph Pros[Advantages]
P1[✅ Fast parsing]
P2[✅ JavaScript native]
P3[✅ Concise syntax]
P4[✅ Widely supported]
end
subgraph Cons[Disadvantages]
C1[❌ No comments]
C2[❌ No trailing commas]
C3[❌ Inconvenient multiline strings]
C4[❌ No date type]
end
JSON --> Pros
JSON --> Cons
JSON Usage Examples
// JavaScript
const config = {
apiUrl: "https://api.example.com",
timeout: 5000,
retries: 3
};
// Convert to JSON
const json = JSON.stringify(config, null, 2);
console.log(json);
// Parse JSON
const parsed = JSON.parse(json);
console.log(parsed.apiUrl);
# Python
import json
config = {
"apiUrl": "https://api.example.com",
"timeout": 5000,
"retries": 3
}
# Convert to JSON
json_str = json.dumps(config, indent=2)
print(json_str)
# Parse JSON
parsed = json.loads(json_str)
print(parsed['apiUrl'])
# File I/O
with open('config.json', 'w') as f:
json.dump(config, f, indent=2)
with open('config.json', 'r') as f:
loaded = json.load(f)
JSON’s strictness is its main strength and its main annoyance. A single trailing comma or a comment makes the whole file invalid, and the error message points at a character offset: Python reports json.decoder.JSONDecodeError: Expecting property name enclosed in double quotes: line 5 column 1, JavaScript reports SyntaxError: Unexpected token } in JSON at position 87 (newer engines phrase it differently). In return, there is exactly one way to read a JSON document, which is why it is the default for data exchange.
Two subtler limits matter for data rather than config. Numbers have no defined precision, and JavaScript parses them as 64-bit floats, so an ID like 12345678901234567890 silently comes back as 12345678901234567000; APIs send large IDs as strings for this reason. And the spec does not say what duplicate keys mean: most parsers keep the last value without warning, so {"debug": false, ..., "debug": true} enables debugging with no error.
JSON Schema (Validation)
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"properties": {
"name": {
"type": "string",
"minLength": 1,
"maxLength": 100
},
"age": {
"type": "integer",
"minimum": 0,
"maximum": 150
},
"email": {
"type": "string",
"format": "email"
},
"tags": {
"type": "array",
"items": {
"type": "string"
},
"minItems": 1
}
},
"required": ["name", "email"]
}
YAML (YAML Ain’t Markup Language)
What is YAML?
YAML is a human-readable data serialization format. It expresses hierarchy through indentation and supports comments.
YAML Basic Syntax
# Comments start with #
# Key-value pairs
name: John Doe
age: 30
email: [email protected]
# Boolean
isActive: true
isDeleted: false
# Null
metadata: null
# or
metadata: ~
# Strings (quotes optional)
city: Seoul
country: "South Korea"
description: 'Single quotes.' # Multiline strings
bio: |
This is a multi-line string.
It preserves line breaks.
Very useful for long text.
summary: >
This is a folded string.
Line breaks are converted to spaces.
Useful for long paragraphs.
# Arrays (lists)
tags:
- developer
- designer
- writer
# Or inline
tags: ["developer", "designer", "writer"]
# Objects (dictionaries)
address:
street: 123 Main St
city: Seoul
zipCode: "12345"
# Or inline
address: {street: 123 Main St, city: Seoul, zipCode: "12345"}
# Arrays + Objects
projects:
- id: 1
name: Project A
status: active
- id: 2
name: Project B
status: completed
# Anchors and aliases (reuse)
defaults: &defaults
timeout: 30
retries: 3
production:
<<: *defaults
apiUrl: https://api.example.com
development:
<<: *defaults
apiUrl: http://localhost:3000
YAML Cautions
# ❌ No tabs (spaces only)
parent:
nested: error # Error: this line is indented with a tab
# ✅ Use spaces (and a key with children has no value on its own line)
parent:
nested: correct
# ❌ Inconsistent indentation
items:
- name: item1
value: 100
- name: item2 # Indentation error!
# ✅ Consistent indentation
items:
- name: item1
value: 100
- name: item2
value: 200
# Special character caution
# A colon followed by a space inside a value needs quotes
url: https://example.com # OK: no space after the colons
title: "Note: read this" # quotes required (": " would start a mapping)
time: "12:30" # Quotes required (YAML 1.1 parsers read 12:30 as the number 750)
# Strings starting with numbers
version: "1.0" # Quotes needed (otherwise parsed as number)
The error messages for these are rarely helpful. A tab produces something like found character '\t' that cannot start any token, and a value with an unquoted ": " produces mapping values are not allowed here, both reported at a line that may be one away from the real mistake.
The deeper problem is implicit typing: YAML decides the type of an unquoted value by what it looks like. PyYAML and many other widely used parsers still implement YAML 1.1, where yes, no, on and off are booleans, so a list of country codes containing NO (Norway) becomes false, and version: 1.10 becomes the float 1.1. YAML 1.2 narrowed booleans to true/false, but you rarely control which version the consuming tool uses. I quote every value that is meant to be a string (versions, IDs, codes, times, anything with a leading zero like "0755") instead of trying to remember which ones are safe; it costs two characters and removes a whole class of bugs.
YAML Usage Examples
# Python
import yaml
config = {
'name': 'MyApp',
'version': '1.0.0',
'database': {
'host': 'localhost',
'port': 5432,
'name': 'mydb'
},
'features': ['auth', 'api', 'admin']
}
# Convert to YAML
yaml_str = yaml.dump(config, default_flow_style=False)
print(yaml_str)
# Parse YAML
parsed = yaml.safe_load(yaml_str)
print(parsed['database']['host'])
# File I/O
with open('config.yaml', 'w') as f:
yaml.dump(config, f, default_flow_style=False)
with open('config.yaml', 'r') as f:
loaded = yaml.safe_load(f)
Always use yaml.safe_load, never plain yaml.load on files you did not write. The full loader can construct arbitrary Python objects from tags such as !!python/object/apply:os.system, which turns a config file into code execution. Since PyYAML 5.1, calling yaml.load without a Loader argument emits a warning, and in PyYAML 6 it is an error (load() missing 1 required positional argument: 'Loader'). Also note that yaml.dump sorts keys alphabetically by default; pass sort_keys=False if the order of a generated file matters to the people reading it, and yaml.dump does not preserve comments at all. Round-tripping a hand-written YAML file through PyYAML deletes every comment; ruamel.yaml exists for that use case.
Docker Compose Example
version: '3.8'
services:
web:
image: nginx:latest
ports:
- "80:80"
- "443:443"
volumes:
- ./html:/usr/share/nginx/html:ro
- ./nginx.conf:/etc/nginx/nginx.conf:ro
environment:
- NGINX_HOST=example.com
- NGINX_PORT=80
depends_on:
- app
networks:
- webnet
restart: unless-stopped
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost/health"]
interval: 30s
timeout: 10s
retries: 3
app:
build:
context: .
dockerfile: Dockerfile
args:
- NODE_ENV=production
ports:
- "3000:3000"
environment:
- DATABASE_URL=postgresql://user:pass@db:5432/mydb
- REDIS_URL=redis://redis:6379
volumes:
- ./app:/app
- /app/node_modules
networks:
- webnet
deploy:
replicas: 3
resources:
limits:
cpus: '0.5'
memory: 512M
db:
image: postgres:15
environment:
POSTGRES_DB: mydb
POSTGRES_USER: user
POSTGRES_PASSWORD: pass
volumes:
- db-data:/var/lib/postgresql/data
networks:
- webnet
networks:
webnet:
driver: bridge
volumes:
db-data:
Two things in this file show how a format’s rules interact with a tool’s rules. The top-level version: '3.8' is obsolete: current Docker Compose ignores it and prints a warning that the attribute version is obsolete, so new files should drop it. And deploy.replicas: 3 combined with ports: - "3000:3000" cannot work on a single host, because three containers cannot bind the same host port; the second and third replicas fail to start with a “port is already allocated” error. YAML accepts both happily; only the tool reading it knows they are wrong, which is why editor schema validation (the Compose and Kubernetes JSON Schemas in the YAML language server) catches more than a YAML linter can.
XML (eXtensible Markup Language)
What is XML?
XML is an extensible markup language that structures data based on tags. It supports strict syntax and schema validation.
XML Basic Syntax
<?xml version="1.0" encoding="UTF-8"?>
<!-- Comment -->
<configuration>
<!-- Simple elements -->
<name>MyApp</name>
<version>1.0.0</version>
<!-- Attributes -->
<database type="postgresql" ssl="true">
<host>localhost</host>
<port>5432</port>
<name>mydb</name>
<credentials>
<username>user</username>
<password>pass</password>
</credentials>
</database>
<!-- Arrays (repeated elements) -->
<features>
<feature name="auth" enabled="true"/>
<feature name="api" enabled="true"/>
<feature name="admin" enabled="false"/>
</features>
<!-- CDATA (includes special characters) -->
<description><![CDATA[
This is a <description> with special characters: & < > " '
]]></description>
<!-- Namespaces -->
<config xmlns="http://example.com/config"
xmlns:db="http://example.com/database">
<db:connection>localhost</db:connection>
</config>
</configuration>
XML Parsing
# Python (ElementTree)
import xml.etree.ElementTree as ET
# Parse XML
tree = ET.parse('config.xml')
root = tree.getroot()
# Access elements
name = root.find('name').text
print(f"Name: {name}")
# Access attributes
db = root.find('database')
db_type = db.get('type')
print(f"Database type: {db_type}")
# Repeated elements
for feature in root.findall('.//feature'):
name = feature.get('name')
enabled = feature.get('enabled')
print(f"Feature {name}: {enabled}")
# Using XPath
host = root.find('.//database/host').text
print(f"Database host: {host}")
Two caveats with this code. find returns None when the element is missing, so root.find('name').text fails with AttributeError: 'NoneType' object has no attribute 'text' rather than a clear “missing key” error. And the namespaced <config xmlns="..."> element in the sample cannot be found as root.find('config'): once a default namespace applies, ElementTree names the element {http://example.com/config}config, and every lookup must include the namespace (or a namespace map). This catches almost everyone the first time they parse Maven POMs or SVG files. For files from untrusted sources, use defusedxml, because standard XML parsers can be abused with entity expansion (“billion laughs”) and, in some parsers, external entity (XXE) attacks.
// JavaScript (Browser)
const parser = new DOMParser();
const xmlDoc = parser.parseFromString(xmlString, "text/xml");
// Access elements
const name = xmlDoc.getElementsByTagName("name")[0].textContent;
console.log(name);
// Access attributes
const db = xmlDoc.getElementsByTagName("database")[0];
const dbType = db.getAttribute("type");
console.log(dbType);
XML Schema (XSD)
<?xml version="1.0" encoding="UTF-8"?>
<xs:schema xmlns:xs="http://www.w3.org/2001/XMLSchema">
<xs:element name="configuration">
<xs:complexType>
<xs:sequence>
<xs:element name="name" type="xs:string"/>
<xs:element name="version" type="xs:string"/>
<xs:element name="database" type="DatabaseType"/>
</xs:sequence>
</xs:complexType>
</xs:element>
<xs:complexType name="DatabaseType">
<xs:sequence>
<xs:element name="host" type="xs:string"/>
<xs:element name="port" type="xs:integer"/>
</xs:sequence>
<xs:attribute name="type" type="xs:string" use="required"/>
</xs:complexType>
</xs:schema>
XML Use Cases
<!-- Maven (pom.xml) -->
<project xmlns="http://maven.apache.org/POM/4.0.0">
<modelVersion>4.0.0</modelVersion>
<groupId>com.example</groupId>
<artifactId>myapp</artifactId>
<version>1.0.0</version>
<dependencies>
<dependency>
<groupId>org.springframework.boot</groupId>
<artifactId>spring-boot-starter-web</artifactId>
<version>3.2.0</version>
</dependency>
</dependencies>
</project>
<!-- Spring (applicationContext.xml) -->
<beans xmlns="http://www.springframework.org/schema/beans">
<bean id="dataSource" class="org.apache.commons.dbcp.BasicDataSource">
<property name="driverClassName" value="org.postgresql.Driver"/>
<property name="url" value="jdbc:postgresql://localhost:5432/mydb"/>
</bean>
</beans>
<!-- Android (AndroidManifest.xml) -->
<manifest xmlns:android="http://schemas.android.com/apk/res/android"
package="com.example.myapp">
<uses-permission android:name="android.permission.INTERNET"/>
<application android:label="MyApp">
<activity android:name=".MainActivity">
<intent-filter>
<action android:name="android.intent.action.MAIN"/>
</intent-filter>
</activity>
</application>
</manifest>
TOML (Tom’s Obvious, Minimal Language)
What is TOML?
TOML is an easy-to-read and clear configuration file format. It’s an improved version of INI, supporting types and nesting.
TOML Basic Syntax
# Comment
# Key-value pairs
name = "MyApp"
version = "1.0.0"
# Numbers
port = 8080
timeout = 30.5
# Boolean
debug = true
production = false
# Date/time
created_at = 2026-04-01T10:00:00Z
# Arrays
tags = ["rust", "web", "api"]
# Inline tables
author = { name = "John Doe", email = "[email protected]" }
# Tables (sections)
[database]
host = "localhost"
port = 5432
name = "mydb"
[database.credentials]
username = "user"
password = "pass"
# Array of tables
[[servers]]
name = "server1"
ip = "192.168.1.1"
role = "primary"
[[servers]]
name = "server2"
ip = "192.168.1.2"
role = "backup"
# Nested structure
[app.cache]
enabled = true
ttl = 3600
[app.cache.redis]
host = "localhost"
port = 6379
TOML avoids YAML’s guessing: strings are always quoted, so version = "1.10" stays a string and version = 1.10 is visibly a float. Dates are a real type, not strings. The trade-off appears with deep nesting. Every level needs a full dotted table header ([app.cache.redis]), so a structure that is five levels deep and has many siblings becomes long and repetitive, which is why Kubernetes-style manifests would be painful in TOML. The other rule people hit is that a table cannot be defined twice: writing [database] in two places is an error (Cannot declare table 'database' twice or similar, depending on the parser), and after [database.credentials] you cannot go back and add keys to [database] by repeating its header.
# Python
import tomli # Python 3.11+: use the built-in tomllib (same API)
with open('config.toml', 'rb') as f: # binary mode is required
config = tomli.load(f)
print(config['name'])
print(config['database']['host'])
print(config['servers'][0]['name'])
# Generate TOML
import tomli_w
data = {
'name': 'MyApp',
'version': '1.0.0',
'database': {
'host': 'localhost',
'port': 5432
}
}
with open('output.toml', 'wb') as f:
tomli_w.dump(data, f)
// Rust
use serde::Deserialize;
use std::fs;
#[derive(Deserialize)]
struct Config {
name: String,
version: String,
database: Database,
}
#[derive(Deserialize)]
struct Database {
host: String,
port: u16,
}
fn main() {
let contents = fs::read_to_string("config.toml").unwrap();
let config: Config = toml::from_str(&contents).unwrap();
println!("Name: {}", config.name);
println!("DB Host: {}", config.database.host);
}
TOML Use Cases
# Cargo.toml (Rust)
[package]
name = "myapp"
version = "0.1.0"
edition = "2021"
[dependencies]
tokio = { version = "1.35", features = ["full"] }
serde = { version = "1.0", features = ["derive"] }
serde_json = "1.0"
[dev-dependencies]
criterion = "0.5"
[profile.release]
opt-level = 3
lto = true
# pyproject.toml (Python)
[project]
name = "myapp"
version = "1.0.0"
description = "My Python application"
authors = [{name = "John Doe", email = "[email protected]"}]
dependencies = [
"fastapi>=0.104.0",
"uvicorn>=0.24.0",
]
[project.optional-dependencies]
dev = ["pytest>=7.4.0", "black>=23.0.0"]
[tool.black]
line-length = 88
target-version = ['py311']
INI (Initialization File)
What is INI?
INI is a simple configuration file format widely used in Windows. It consists of sections and key-value pairs.
INI Basic Syntax
; Comments start with semicolon
# Or hash (#) is also possible
; Global settings: Python's configparser requires a section header
[DEFAULT]
app_name = MyApp
version = 1.0.0
; Section
[database]
host = localhost
port = 5432
name = mydb
username = user
password = pass
[cache]
enabled = true
ttl = 3600
type = redis
[cache.redis]
host = localhost
port = 6379
; Arrays (non-standard, varies by implementation)
[features]
feature1 = auth
feature2 = api
feature3 = admin
; Or
features = auth,api,admin
INI Parsing
# Python
import configparser
config = configparser.ConfigParser()
config.read('config.ini')
# Read values
app_name = config['DEFAULT']['app_name']
db_host = config['database']['host']
db_port = config.getint('database', 'port')
cache_enabled = config.getboolean('cache', 'enabled')
print(f"App: {app_name}")
print(f"DB: {db_host}:{db_port}")
print(f"Cache: {cache_enabled}")
# Write values
config['database']['host'] = 'db.example.com'
with open('config.ini', 'w') as f:
config.write(f)
INI has no specification, and each parser fills the gaps differently, which is the real cost of its simplicity. Python’s configparser shows several of the differences. Keys before the first section header are rejected with configparser.MissingSectionHeaderError: File contains no section headers, which is why the sample puts global settings under [DEFAULT]; values in [DEFAULT] are then inherited by every other section. Every value is a string until you ask for getint/getboolean. Keys are lowercased by default, so ApiKey and apikey collide. Inline comments are not stripped (host = localhost ; dev gives the value localhost ; dev) unless you pass inline_comment_prefixes. And % starts an interpolation, so a password containing % raises InterpolationSyntaxError unless you use RawConfigParser or write %%. [cache.redis] is just a section whose name contains a dot; unlike TOML, it is not nested under [cache].
INI Use Cases
; php.ini
[PHP]
engine = On
short_open_tag = Off
precision = 14
output_buffering = 4096
zlib.output_compression = Off
implicit_flush = Off
serialize_precision = -1
disable_functions = exec,passthru,shell_exec,system
[Date]
date.timezone = Asia/Seoul
[Session]
session.save_handler = files
session.save_path = "/var/lib/php/sessions"
session.gc_maxlifetime = 1440
; .gitconfig
[user]
name = John Doe
email = [email protected]
[core]
editor = vim
autocrlf = input
[alias]
st = status
co = checkout
br = branch
ci = commit
Markdown
What is Markdown?
Markdown is a lightweight markup language for writing formatted documents in plain text.
Markdown Basic Syntax
# Heading 1 (H1)
## Heading 2 (H2)
### Heading 3 (H3)
**Bold** or __Bold__
*Italic* or _Italic_
~~Strikethrough~~
`Inline code`
> Blockquote
> Multiple lines possible
- Unordered list
- Item 2
- Nested item
- Nested item 2
1. Ordered list
2. Item 2
3. Item 3
[Link text](https://example.com)

---
Horizontal rule
| Column 1 | Column 2 | Column 3 |
|----------|----------|----------|
| Value 1 | Value 2 | Value 3 |
| Value 4 | Value 5 | Value 6 |
```code block```
Multiple lines of code
```python
# Language specification (syntax highlighting)
def hello():
print("Hello, World!")
```
- [ ] Checkbox (incomplete)
- [x] Checkbox (complete)
GitHub Flavored Markdown (GFM)
# GitHub Extensions
## Task Lists
- [x] Completed task
- [ ] Incomplete task
- [ ] In progress
## Table Alignment
| Left | Center | Right |
|:-----|:------:|------:|
| Left | Center | Right |
## Emoji
:smile: :rocket: :tada:
## Mentions
@username
## Issue References
#123
## Code Block with Language (syntax highlighting)
```javascript
console.log("Hello");
```
## Collapsible Sections
<details>
<summary>Click to expand</summary>
Hidden content
</details>
## Math (LaTeX)
$E = mc^2$
$$
\frac{-b \pm \sqrt{b^2 - 4ac}}{2a}
$$
Markdown Parsing
# Python (markdown library)
import markdown
md_text = """
# Hello
This is **bold** and *italic*.
- Item 1
- Item 2
"""
html = markdown.markdown(md_text)
print(html)
# <h1>Hello</h1>
# <p>This is <strong>bold</strong> and <em>italic</em>.</p>
# <ul>
# <li>Item 1</li>
# <li>Item 2</li>
# </ul>
# Using extensions
html = markdown.markdown(md_text, extensions=['tables', 'fenced_code', 'toc'])
// JavaScript (marked library)
import { marked } from 'marked';
const mdText = `
# Hello
This is **bold** text.
`;
const html = marked.parse(mdText);
console.log(html);
7. Format Comparison and Selection Guide
Same Data Representation
JSON
{
"app": {
"name": "MyApp",
"version": "1.0.0",
"features": ["auth", "api"]
},
"database": {
"host": "localhost",
"port": 5432
}
}
YAML
# Example
app:
name: MyApp
version: 1.0.0
features:
- auth
- api
database:
host: localhost
port: 5432
XML
<?xml version="1.0"?>
<config>
<app>
<name>MyApp</name>
<version>1.0.0</version>
<features>
<feature>auth</feature>
<feature>api</feature>
</features>
</app>
<database>
<host>localhost</host>
<port>5432</port>
</database>
</config>
TOML
# Example
[app]
name = "MyApp"
version = "1.0.0"
features = ["auth", "api"]
[database]
host = "localhost"
port = 5432
INI
[app]
name = MyApp
version = 1.0.0
features = auth,api
[database]
host = localhost
port = 5432
Comparison Table
| Feature | JSON | YAML | XML | TOML | INI | Markdown |
|---|---|---|---|---|---|---|
| Readability | ⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ |
| Parsing Speed | ⭐⭐⭐⭐⭐ | ⭐⭐ | ⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐ |
| Comments | ❌ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Type Safety | ⭐⭐⭐ | ⭐⭐ | ⭐⭐⭐⭐⭐ | ⭐⭐⭐⭐ | ⭐⭐ | N/A |
| Nested Structure | ✅ | ✅ | ✅ | ✅ | Limited | ✅ |
| File Size | Small | Medium | Large | Medium | Small | Medium |
| Schema Validation | ✅ | ✅ | ✅ | ❌ | ❌ | ❌ |
Recommendations by Use Case
flowchart TD
Start[Choose File Format] --> Q1{Purpose?}
Q1 -->|API Response| JSON["✅ JSON\nFast parsing"]
Q1 -->|Config File| Q2{Complexity?}
Q1 -->|Documentation| Markdown["✅ Markdown\nREADME, blog"]
Q1 -->|Build Config| Q3{Language?}
Q2 -->|Simple| INI["✅ INI\nSimple config"]
Q2 -->|Medium| TOML["✅ TOML\nRust, Python"]
Q2 -->|Complex| YAML["✅ YAML\nDocker, K8s"]
Q3 -->|Java| XML["✅ XML\nMaven, Spring"]
Q3 -->|JavaScript| JSON["✅ JSON\npackage.json"]
Q3 -->|Rust| TOML["✅ TOML\nCargo.toml"]
Q3 -->|Python| TOML2["✅ TOML\npyproject.toml"]
Project Usage
| Project | Format | Filename |
|---|---|---|
| Node.js | JSON | package.json |
| Python | TOML | pyproject.toml |
| Rust | TOML | Cargo.toml |
| Go | Go | go.mod |
| Docker | YAML | docker-compose.yml |
| Kubernetes | YAML | deployment.yaml |
| Ansible | YAML | playbook.yml |
| Maven | XML | pom.xml |
| Gradle | Groovy | build.gradle |
| Nginx | Custom | nginx.conf |
| Apache | Custom | httpd.conf |
8. Practical Conversion and Validation
Format Conversion
#!/usr/bin/env python3
"""
Configuration file format conversion tool
"""
import json
import yaml
import tomli
import tomli_w
import xml.etree.ElementTree as ET
from pathlib import Path
class ConfigConverter:
@staticmethod
def json_to_yaml(json_file, yaml_file):
"""JSON → YAML"""
with open(json_file, 'r') as f:
data = json.load(f)
with open(yaml_file, 'w') as f:
yaml.dump(data, f, default_flow_style=False, allow_unicode=True)
print(f"✅ Converted: {json_file} → {yaml_file}")
@staticmethod
def yaml_to_json(yaml_file, json_file):
"""YAML → JSON"""
with open(yaml_file, 'r') as f:
data = yaml.safe_load(f)
with open(json_file, 'w') as f:
json.dump(data, f, indent=2, ensure_ascii=False)
print(f"✅ Converted: {yaml_file} → {json_file}")
@staticmethod
def json_to_toml(json_file, toml_file):
"""JSON → TOML"""
with open(json_file, 'r') as f:
data = json.load(f)
with open(toml_file, 'wb') as f:
tomli_w.dump(data, f)
print(f"✅ Converted: {json_file} → {toml_file}")
@staticmethod
def toml_to_json(toml_file, json_file):
"""TOML → JSON"""
with open(toml_file, 'rb') as f:
data = tomli.load(f)
with open(json_file, 'w') as f:
json.dump(data, f, indent=2, ensure_ascii=False)
print(f"✅ Converted: {toml_file} → {json_file}")
# Usage
converter = ConfigConverter()
converter.json_to_yaml('config.json', 'config.yaml')
converter.yaml_to_json('config.yaml', 'config.json')
converter.json_to_toml('config.json', 'config.toml')
Conversion is lossless only in the direction from stricter to looser formats, and even then with caveats. JSON → YAML always works, since every JSON document is valid YAML 1.2. YAML → JSON loses comments and anchors (aliases are expanded into copies), and converts dates to strings. JSON → TOML fails outright for data TOML cannot represent: null has no TOML equivalent, so tomli_w.dump raises a TypeError on any None value, and a top-level array is not allowed because a TOML document is always a table. Converting a real config is therefore a one-time migration you check by hand, not a round trip you can automate safely.
Validation Tools
# JSON validation
jq . config.json
# or
python -m json.tool config.json
# YAML validation
yamllint config.yaml
# YAML → JSON (yq)
yq eval -o=json config.yaml
# XML validation
xmllint --noout config.xml
# TOML validation (Python)
python -c "import tomli; tomli.load(open('config.toml', 'rb'))"
Online Tools
# jq (JSON query)
cat config.json | jq '.database.host'
# "localhost"
# Array filtering
cat config.json | jq '.servers[] | select(.role == "primary")'
# Conversion
cat config.json | jq '.database'
# yq (YAML query)
cat config.yaml | yq '.database.host'
# xmllint (XML query)
xmllint --xpath '//database/host/text()' config.xml
Other Formats
Properties (Java)
# application.properties (Spring Boot)
server.port=8080
server.address=0.0.0.0
spring.datasource.url=jdbc:postgresql://localhost:5432/mydb
spring.datasource.username=user
spring.datasource.password=pass
# Arrays
spring.profiles.active=dev,debug
# Multiline (backslash)
app.description=This is a very long \
description that spans \
multiple lines.
HCL (HashiCorp Configuration Language)
# Terraform
variable "region" {
description = "AWS region"
type = string
default = "us-west-2"
}
resource "aws_instance" "web" {
ami = "ami-0c55b159cbfafe1f0"
instance_type = "t2.micro"
tags = {
Name = "WebServer"
Environment = "production"
}
}
output "instance_ip" {
value = aws_instance.web.public_ip
}
Jsonnet (JSON Template)
// config.jsonnet
local env = std.extVar('env');
local base = {
name: 'MyApp',
version: '1.0.0',
};
local envConfig = {
dev: {
apiUrl: 'http://localhost:3000',
debug: true,
},
prod: {
apiUrl: 'https://api.example.com',
debug: false,
},
};
base + envConfig[env]
# Execute
jsonnet -V env=dev config.jsonnet
# {
# "name": "MyApp",
# "version": "1.0.0",
# "apiUrl": "http://localhost:3000",
# "debug": true
# }
ENV Files
# .env
NODE_ENV=production
PORT=3000
DATABASE_URL=postgresql://user:pass@localhost:5432/mydb
REDIS_URL=redis://localhost:6379
API_KEY=secret-key-123
DEBUG=false
# Arrays (non-standard)
ALLOWED_HOSTS=localhost,example.com,*.example.com
# Python (python-dotenv)
from dotenv import load_dotenv
import os
load_dotenv()
port = int(os.getenv('PORT', 3000))
database_url = os.getenv('DATABASE_URL')
debug = os.getenv('DEBUG', 'false').lower() == 'true'
print(f"Port: {port}")
print(f"Database: {database_url}")
print(f"Debug: {debug}")
Comments, per-environment values, secrets and validation
Comments in JSON-like files
// JSON5: comments, unquoted keys and trailing commas
// (only works where the reading tool supports JSON5)
{
// Project info
name: "myapp",
version: "1.0.0",
// Dependencies
dependencies: {
express: "^4.18.0", // Trailing commas allowed
},
}
// Plain JSON workaround: a _comment key (ignored only by convention)
{
"_comment": "This is a comment",
"name": "myapp"
}
JSON5 and JSONC only help when the tool reading the file supports them. npm reads only package.json and will not pick up a package.json5, and a JSON5 file passed to JSON.parse fails immediately. VS Code settings and tsconfig.json accept JSONC (comments and trailing commas), which is why comments work there and nowhere else. The _comment key is not free either: a strict schema with additionalProperties: false rejects it. If you need comments in a config you design yourself, that is usually a sign the file should be YAML or TOML.
Per-environment sections with YAML anchors
# config.yaml
default: &default
timeout: 30
retries: 3
development:
<<: *default
apiUrl: http://localhost:3000
debug: true
production:
<<: *default
apiUrl: https://api.example.com
debug: false
# Load by environment with Python
import yaml
import os
env = os.getenv('ENV', 'development')
with open('config.yaml') as f:
config = yaml.safe_load(f)
app_config = config[env]
The << merge key comes from YAML 1.1 and is not part of the YAML 1.2 core schema. PyYAML and most YAML libraries still honour it, but a strict 1.2 parser may treat << as an ordinary key, so check before relying on it in a file that several tools read. The merge is also shallow: if default contains a nested mapping and production redefines that mapping, the whole nested mapping is replaced rather than merged key by key.
Keeping secrets out of the file
# config.yaml (exclude sensitive info)
database:
host: localhost
port: 5432
name: mydb
# username and password from environment variables
# .env (exclude from version control)
DB_USERNAME=user
DB_PASSWORD=secret
# .gitignore
.env
*.secret.yaml
# Merge with Python
import yaml
import os
with open('config.yaml') as f:
config = yaml.safe_load(f)
# Override with environment variables
config['database']['username'] = os.getenv('DB_USERNAME')
config['database']['password'] = os.getenv('DB_PASSWORD')
os.getenv returns None when the variable is missing, so this code starts happily and fails later at connection time. Use os.environ['DB_PASSWORD'] (which raises KeyError) or the schema check below if a missing secret should stop startup.
Validating the parsed result
# JSON Schema validation
import json
import jsonschema
schema = {
"type": "object",
"properties": {
"name": {"type": "string"},
"age": {"type": "integer", "minimum": 0}
},
"required": ["name"]
}
data = {"name": "John", "age": 30}
try:
jsonschema.validate(data, schema)
except jsonschema.ValidationError as e:
print(f"Invalid config: {e.message}")
JSON Schema validates data, not a file format, so the same schema works for a YAML or TOML file after it is parsed into dictionaries. That is also where YAML’s type guessing gets caught: a version: 1.10 that became the float 1.1 fails a "type": "string" rule at startup.
Choosing a format by who writes the file
| Format | Best Advantage | Biggest Disadvantage | Recommended Use |
|---|---|---|---|
| JSON | Fast parsing | No comments | API, data exchange |
| YAML | Readability | Slow parsing | Docker, K8s, CI/CD |
| XML | Schema validation | Verbose | Java, SOAP, RSS |
| TOML | Clear syntax | Limited support | Rust, Python config |
| INI | Simplicity | No standard | Simple config |
| Markdown | Easy to read | Limited data structure | Docs, README |
The most useful question is who writes the file. When programs write it and read it back (API payloads, lockfiles, generated manifests), use JSON: it is unambiguous, every language parses it the same way, and the lack of comments does not matter. When people edit it by hand, they need comments, so JSON becomes awkward. TOML is the safer choice there for application and tool settings, because its values have explicit types. YAML is often not a choice at all, since Kubernetes, CI systems and Compose require it; when you do use it, quote strings that could be read as numbers, booleans or dates, which is the source of most YAML surprises (see the FAQ below).
Keep XML for ecosystems built around it (Maven, Android resources, SOAP) and INI for software that already expects it. Whatever you pick, keep secrets out of the file itself and validate the parsed result against a schema at startup, so a typo fails immediately instead of at the first request that reads the setting.
Conversion Cheat Sheet
Command Line Tools
# JSON → YAML
cat config.json | yq -P > config.yaml
# YAML → JSON
cat config.yaml | yq -o=json > config.json # mikefarah/yq (Go); the Python yq wrapper uses different flags
# JSON formatting
cat config.json | jq . > formatted.json
# YAML validation
yamllint config.yaml
# XML formatting
xmllint --format config.xml
# JSON merge
jq -s '.[0] * .[1]' base.json override.json > merged.json
# YAML merge
yq eval-all 'select(fileIndex == 0) * select(fileIndex == 1)' base.yaml override.yaml
By Programming Language
# Python: All formats supported
import json # Built-in
import yaml # pip install pyyaml
import tomllib # Built-in since 3.11 (pip install tomli before that)
import configparser # Built-in (INI)
import xml.etree.ElementTree as ET # Built-in
# JavaScript/Node.js
const json = require('./config.json'); // Native
const yaml = require('js-yaml');
const toml = require('toml');
const ini = require('ini');
# Go
import (
"encoding/json"
"gopkg.in/yaml.v3"
"github.com/BurntSushi/toml"
"gopkg.in/ini.v1"
)
# Rust
use serde_json; // JSON
use serde_yaml; // YAML
use toml; // TOML
Frequently Asked Questions (FAQ)
Q. Why does YAML turn my version number or time value into something else?
A. YAML infers types from unquoted values, so version: 1.10 becomes the float 1.1 and values such as 12:30 can be parsed differently by different parsers. In YAML 1.1 parsers, words like yes, no, on and off become booleans, which is how a country code like NO turns into false. Quote any value that must stay a string, as the YAML Cautions section recommends. TOML and JSON avoid this because their string, number and boolean syntax is explicit.
Related Articles (Internal Links)
Other articles related to this topic.
- Docker Compose
- .env Files and dotenv
- Why Text Gets Garbled: ASCII, Unicode, UTF-8 and BOM
- Building Static Sites with Astro