How to Extract Texture Paths from JSON and Audit Local Directories with Python
28 Aug 26 (2mo ago)
I was recently handed a bunch of project files to render. I opened them up and quickly ran into the classic problem: dozens of textures were linked in the scene, but half of them were nowhere to be found in the asset folders sent over.
Manually cross-referencing shader trees to see what's actually missing is miserable. Some render engines let you view a texture manager or grab a screenshot of missing links, but when you’re dealing with multiple shots or messy material trees, it gets tedious fast.
To stop wasting time, I wrote this script. It parses an exported JSON file of the material/shader network, grabs every linked texture file, and compares it directly against the local tex directory. If anything is missing—or if there's unused bloat—it spits out a clean diagnostic report I can send right back to the client.
Handling Nested Material Trees
Exported material data (like V-Ray or other node-based shader trees) is almost never flat. Maps are buried deep inside sub-materials, layered shaders, mix nodes, and utility slots. On top of that, DCC exports often prepend engine-specific junk like Bitmap: or HDRI: to the path string.
To handle that cleanly, the script needs to:
- Walk recursively through every level of the JSON tree.
- Strip out application prefixes to isolate the actual file path.
- Compare just the filenames against what's physically on disk using Python sets.
The Script
import json
import os
# --- CONFIGURATION ---
JSON_FILE = 'scene_materials.json'
TARGET_DIR = r'D:\girl\c4d\tex' # Path to your local texture directory
def extract_textures(data, found_textures=None):
"""Recursively traverses JSON to find 'value' keys holding file paths."""
if found_textures is None:
found_textures = []
if isinstance(data, dict):
for key, value in data.items():
if key == "value" and isinstance(value, str):
# Filter out unassigned slots and empty strings
if "None" not in value and value.strip():
# Strip out common engine prefixes
clean_path = value.replace("HDRI: ", "").replace("Bitmap: ", "").strip()
found_textures.append(clean_path)
elif isinstance(value, (dict, list)):
extract_textures(value, found_textures)
elif isinstance(data, list):
for item in data:
extract_textures(item, found_textures)
return found_textures
def run_diagnostics():
# 1. Load JSON and extract referenced paths
try:
with open(JSON_FILE, 'r', encoding='utf-8') as f:
parsed_json = json.load(f)
raw_texture_list = extract_textures(parsed_json)
except FileNotFoundError:
print(f"CRITICAL ERROR: Could not find '{JSON_FILE}'. Check your file path.")
return
# 2. Verify target directory exists
if not os.path.exists(TARGET_DIR):
print(f"CRITICAL ERROR: Directory '{TARGET_DIR}' does not exist.")
return
# 3. Normalize filenames for clean set comparison
# Normalize backslashes to handle cross-platform path exports
json_filenames = {
os.path.basename(p.replace("\\", "/"))
for p in raw_texture_list
if os.path.basename(p.replace("\\", "/"))
}
local_filenames = {
f for f in os.listdir(TARGET_DIR)
if os.path.isfile(os.path.join(TARGET_DIR, f))
}
# 4. Set logic operations
matches = json_filenames.intersection(local_filenames)
missing = json_filenames.difference(local_filenames)
extra = local_filenames.difference(json_filenames)
# 5. Diagnostic Output
print("=" * 60)
print("TEXTURE DIAGNOSTICS REPORT")
print(f"Target Directory: {TARGET_DIR}")
print("=" * 60)
print("\n[ SUMMARY ]")
print(f" • Required in JSON: {len(json_filenames)}")
print(f" • Verified in Folder: {len(matches)}")
print(f" • Missing from Folder: {len(missing)}")
print(f" • Extra / Unused: {len(extra)}")
if matches:
print(f"\n[ MATCHED FILES ] ({len(matches)})")
for f in sorted(matches):
print(f" [OK] {f}")
if missing:
print(f"\n[ MISSING FILES ] ({len(missing)} files need to be requested)")
for f in sorted(missing):
print(f" [MISSING] {f}")
if extra:
print(f"\n[ UNUSED / EXTRA FILES ] ({len(extra)} files in folder not referenced)")
for f in sorted(extra):
print(f" [EXTRA] {f}")
print("\n" + "=" * 60)
if __name__ == "__main__":
run_diagnostics()
How It Works
- Recursive Traversal:
extract_textures()doesn't care how deep a map is nested inside a multi-layered material node. It inspects every dictionary key and list element until it finds valid string values. - Path Cleaning: Strips DCC prefixes (
HDRI:,Bitmap:) and normalizes paths so Windows/Unix slashes don't breakos.path.basename(). - Fast Set Operations: Instead of running slow nested loops across hundreds of files, it converts both lists into Python
set()objects.difference()andintersection()instantly tell you what’s accounted for, what’s completely missing, and what extra garbage is sitting in your project directory.