Updates May 25
PGrid
- running evals for classification and specs search
- deepseek v4 pro is worse than gpt-5.4-mini at classification
- codex gpt-5.4-mini seems pretty good for search grounded specs search prompts, but still not up to gemini flash grounded quality
- added specs search prompt evals and dual-client consensus mode for review replay evals
- added make command to bootstrap libraries and pnpm on remote vps. this needed a bit of iterating as gpt-5.5 kept getting it wrong
- work in progress on officeworks crawler + laptop support
- did updates to officeworks crawler and crawler skill
- enabled officeworks for laptops
- added support for single item classification
- supports cpu/gpu/ssd/ram. this is to allow specific items to fast track through classification
- added a cli command for this
- updated laptop variant specs search to include db examples for same model
- updated eval and logging tooling
- added eval logging updates and made logging more consistent across categories
- passed reasoning effort through for local codex