SIGGRAPH-2022 CLIPasso:Semantically-Aware Object Sketching
Paper: CLIPasso:Semantically-Aware Object Sketching
Project page: https://clipasso.github.io/clipasso/
CLIPasso: object sketches via geometric and semantic losses
Abstract
CLIPasso is introduced as a method that renders objects at different levels of abstraction through geometric and semantic simplification. Because sketch generation traditionally depends heavily on sketch datasets for training, CLIP is used to extract semantic information from both sketches and images.
Introduction


Turning an image into a sketch is very difficult for computers: the system must identify salient features, and abstraction is hard to achieve.

Prior work fixes abstraction in advance—more strokes imply less abstraction, fewer strokes imply more—and models largely reflect whatever sketch data they were trained on. This data-driven pipeline constrains style and form, runs counter to the spirit of image processing, and yields limited diversity.
Escaping supervised sketch data calls for a model with strong semantic understanding; CLIP fits this role. Trained on paired text and images, it captures semantics well, is sensitive to object identity rather than image style, and offers strong zero-shot behavior that can be used directly.
Method

Each stroke is controlled by four points, $s_i={p_i^j}^4j=1={(x_i,y_i)^j}^4{j=1}$. The model adjusts these four control points to define the curve; Bézier evaluation then shapes the curve into the final sketch.
CLIP acts as a teacher for distilling the optimized sketch representation.
Loss Function

If the target is described as a horse, the sketch features and the original image features should both encode “horse,” yielding the semantic loss $L_s$.

Sketch features are pushed to match the original features, but semantic agreement alone is insufficient: two horses can match semantically while the sketch places the head on the wrong side. A geometric term $L_g$ therefore complements the semantic constraint. Features are taken from stages 2–4 of a ResNet-50 backbone rather than the final 2048-D vector; using earlier layers better preserves shape and orientation.

Optimization

Optimization runs for 2000 iterations but is fast; quality is already apparent after roughly 100 iterations.
Stroke Initialization

Initial placement of Bézier control points matters. The authors use saliency-guided initialization: a pretrained vision transformer supplies multi-head self-attention from the last layer (weighted average), and points are sampled in more salient regions.
Result

Sketches can be produced for uncommon objects as well.

Abstraction level is controlled by the number of strokes initialized at the start.

Compared with other methods, CLIPasso achieves stronger abstract sketches.
Limitations
Performance drops when the image contains background; a single foreground object works best. The authors segment the object and sketch the foreground, but this two-stage pipeline may not be optimal.
Human sketching is sequential—one stroke after another—whereas CLIPasso generates all strokes jointly.
Using stroke count to control abstraction is both a strength and a limitation: the user must preset the count, and matching a desired abstraction level may require different stroke numbers that are hard to specify in advance. Letting the model decide whether two or three strokes (or more) suffice would be preferable.
<HR align=left color=#987cb9 SIZE=1>