Each loop iterates: eval the parsing agent → propose a tool/prompt change → A/B eval → adopt or revert — until the quality target is reached.