The paper explores whether long text can be compressed by converting it into an image and feeding that image to a multimodal LLM instead of sending all text tokens directly. The key idea is that visio
The document argues that apparent “human-like” (anthropomorphic) attributes in large language models (LLMs)—such as understanding, morality, empathy, deception, or self-awareness—cannot be treated as