
GPT-6 Astra’s high token consumption stems from its lack of memory, as each message re-processes the entire conversation history, causing costs to compound rapidly. To optimize spending, users should avoid defaulting to "max effort" settings, which can cost four times more than lower tiers for identical tasks. Instead, treat Astra as a senior developer, delegating routine tasks to cheaper models and reserving Astra only for complex judgment calls. Monitoring usage is critical, specifically staying below the 272,000-token threshold where significant pricing multipliers trigger. Additionally, avoid common inefficiencies like using PDFs, which incur double charges, or switching models mid-session, which forces a full-cost re-process. By implementing these strategies—such as using sub-agents for routing and cleaning up tool instructions—users can significantly reduce expenses while maintaining high-quality output.
Sign in to continue reading, translating and more.
Open full episode in Podwise