Context-Dependent Affordance Reports in Vision-Language Models
Separating Persona Effects from Measurement Artifacts
Version note · 19 September 2026
Corrected version published on Zenodo on 12 September 2026 (v3.0.0, DOI 10.5281/zenodo.22721059). The downloadable PDF is that published release. It replaces the earlier empty-response analysis, percentage-of-meaning claims and functional interpretation of the Tucker factor. The arXiv replacement was submitted separately.
Summary
Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify an affordance effect. An audit of the earlier seven-prompt study separates empty responses from content comparisons and withdraws the functional-manifold interpretation and percentages-of-meaning claims. A matched-question experiment uses 48 new images, four personas, a shared three-object task, two wordings and two requested seeds in each of two model configurations: 1,536 requests in total. Persona-associated variation does not uniformly exceed wording or sampling variation. The results describe context-conditioned reports, with explicit limits on inference about internal processing or embodied behaviour.