A GPT-4-based agent asked to complete a realistic task on a live website, book a flight, file a support ticket, or post to a forum succeeds only 14.41% of the time end-to-end, against a human success rate of 78.24%, according to…