arXiv preprint · cs.CV
Image-Space Rule Discovery
Introduces WISRD, a benchmark for testing whether image-editing models can interpret visual instructions, infer rules, and complete worksheet-style tasks directly in image space.
Research profile
杉山 未空
Computer Vision · Image & Video Generation

02 — Publications
Two public arXiv preprints spanning computer vision, generative modeling, image editing, and visual reasoning.
arXiv preprint · cs.CV
Introduces WISRD, a benchmark for testing whether image-editing models can interpret visual instructions, infer rules, and complete worksheet-style tasks directly in image space.
arXiv · ICCV 2025 Workshop
A multi-label framework for detecting four common types of visual artifacts in Sora-generated videos, achieving 94.14% average classification accuracy with ResNet-50.
03 — What's new?
04 — Projects
Learning representations that capture objects, structure, context, and change in visual data.
Studying generative models for coherent, controllable, and expressive visual synthesis.
Translating images across domains while preserving the content and intent that matter.
05 — Bio
My work sits at the intersection of computer vision and generative AI. I am interested in models that do more than create convincing pixels—models that understand structure, preserve intent, and translate visual information across domains.
Current interests include image and video generation, image-to-image models, controllable synthesis, and the representations that connect perception with generation.
06 — Affiliations
Academic affiliation
Research community
Computer vision community
07 — Contact
For research inquiries, talks, or potential collaborations, feel free to get in touch.