Skip to main content

Technical reference

The Brightspot DAM plugin adds download, format conversion, and document data extraction to asset types. It defines the interfaces that asset types implement, the options and format classes that control conversion, and the extraction services that populate text and thumbnails. Optional submodules add converters and extractors that depend on external tools or services.

Dependencies​

  • Pandoc submodule: the pandoc executable on the application server, plus pdflatex or wkhtmltopdf for PDF output
  • ImageMagick submodule: the convert executable on the application server
  • DocRaptor submodule: a DocRaptor account and API key
  • CloudConvert submodule: a CloudConvert account and API key
  • Textract submodule: the com.psddev:aws-textract integration and an Amazon Textract setup with an SQS queue, an SNS topic, and an IAM role. See Amazon Textract configuration.

Installation​

To add the core plugin:

<!-- Requires Brightspot 5.0 or later. -->
<dependency>
<groupId>com.brightspot.dam</groupId>
<artifactId>dam</artifactId>
<version>1.2.0</version>
</dependency>

Add only the submodules that your project needs:

ArtifactAdds
pandocConversion of documents to Word, EPUB, Markdown, PDF, and other formats using Pandoc.
imagemagickConversion of images to formats such as JPG, PNG, TIFF, and GIF using ImageMagick.
docraptorHTML to PDF conversion using DocRaptor.
pdfboxText and thumbnail extraction from PDF files.
cloudconvertText and thumbnail extraction using CloudConvert.
textractText extraction from PDF, JPEG, and PNG files using Amazon Textract.
twelvemonkeysAdditional image formats for image downloads, such as TIFF, PICT, and PNM.
bomA bill of materials that aligns the versions of all DAM modules.
<!-- Requires Brightspot 5.0 or later. -->
<dependency>
<groupId>com.brightspot.dam</groupId>
<artifactId>pandoc</artifactId>
<version>1.2.0</version>
</dependency>

Making assets downloadable​

An asset type is downloadable when it implements Downloadable or one of its subinterfaces. Brightspot discovers the available DownloadOptions classes at run time, so adding a submodule to the build enables its formats without further configuration.

InterfaceModuleUse for
DocumentDownloadablecoreDocuments that Brightspot renders as HTML through a view model.
PandocDownloadablepandocDocuments stored in a StorageItem, such as a Word file.
ImageDownloadablecoreImages in a StorageItem.
ImageMagickDownloadableimagemagickImages that ImageMagick converts.

Documents rendered as HTML​

Implement DocumentDownloadable on the content type. The default DocumentDownloadable#downloadDocumentFiles implementation calls DocumentDownloadOptions#download with no StorageItem, so Brightspot renders the asset through its view model.

1
import com.psddev.cms.db.Content;
2
import com.psddev.dam.DocumentDownloadable;
3
4
public class ReportDocument extends Content implements DocumentDownloadable {
5
6
private String title;
7
8
private String body;
9
10
public String getTitle() {
11
return title;
12
}
13
14
public void setTitle(String title) {
15
this.title = title;
16
}
17
18
public String getBody() {
19
return body;
20
}
21
22
public void setBody(String body) {
23
this.body = body;
24
}
25
}

Then implement the DamDocumentEntryView marker interface on the view model that renders the document. Brightspot renders the view model to HTML, embeds remote images as data URIs, and converts the HTML to the selected format.

1
import com.psddev.cms.view.ViewModel;
2
import com.psddev.dam.view.DamDocumentEntryView;
3
4
public class ReportDocumentViewModel extends ViewModel<ReportDocument> implements DamDocumentEntryView {
5
6
public String getTitle() {
7
return model.getTitle();
8
}
9
10
public String getBody() {
11
return model.getBody();
12
}
13
}

Core formats for this path are text and Markdown. The pandoc module adds more, and the docraptor module adds PDF.

Documents stored in a file​

Implement PandocDownloadable and pass the StorageItem to PandocDocumentDownloadOptions#download. Pandoc deduces the source format from the file extension. To set the source format explicitly, override PandocDownloadable#fromFormat and return a PandocFromFormat value.

1
import java.io.IOException;
2
import java.nio.file.Path;
3
4
import com.psddev.cms.db.Content;
5
import com.psddev.dam.pandoc.PandocDocumentDownloadOptions;
6
import com.psddev.dam.pandoc.PandocDownloadable;
7
import com.psddev.dam.pandoc.PandocFromFormat;
8
import com.psddev.dari.util.StorageItem;
9
10
public class WordDocument extends Content implements PandocDownloadable {
11
12
private StorageItem wordDocumentStorageItem;
13
14
@Override
15
public PandocFromFormat fromFormat() {
16
return PandocFromFormat.WORD_DOCUMENT;
17
}
18
19
@Override
20
public void downloadPandocFiles(PandocDocumentDownloadOptions options, Path root) throws IOException {
21
options.download(this, wordDocumentStorageItem, root);
22
}
23
}

Images​

Implement ImageDownloadable. The default ImageDownloadable#downloadImageFiles implementation uses the preview StorageItem and throws an IOException if the asset has none. Override the method to use a different file.

1
import java.io.IOException;
2
import java.nio.file.Path;
3
4
import com.psddev.cms.db.Content;
5
import com.psddev.dam.ImageDownloadOptions;
6
import com.psddev.dam.ImageDownloadable;
7
import com.psddev.dari.util.StorageItem;
8
9
public class PhotoAsset extends Content implements ImageDownloadable {
10
11
private StorageItem image;
12
13
@Override
14
public void downloadImageFiles(ImageDownloadOptions options, Path root) throws IOException {
15
options.download(this, image, root);
16
}
17
}

With no format selected, DefaultImageFormat keeps the original file extension and falls back to JPG if Java cannot write that format. The core module provides JPG, PNG, GIF, and BMP formats. The imagemagick and twelvemonkeys modules provide more.

Extracting text and thumbnails​

Implement DocumentDataExtractable on the asset type. DocumentDataExtractable#getDocumentFile returns the first file field on the type by default. Override it to return a different StorageItem.

1
import com.psddev.cms.db.Content;
2
import com.psddev.dam.DocumentDataExtractable;
3
import com.psddev.dari.util.StorageItem;
4
5
public class SearchablePdf extends Content implements DocumentDataExtractable {
6
7
private StorageItem file;
8
9
@Override
10
public StorageItem getDocumentFile() {
11
return file;
12
}
13
}

When an asset is saved, DocumentDataExtractableData compares the file with the stored version. If the asset is new or the file changed, Brightspot clears the existing text and thumbnail and submits a DocumentExtractionTask that runs the first configured DocumentDataExtractor whose shouldRun method returns true for the file's content type. The results are saved in the text and thumbnail fields of DocumentDataExtractableData. The thumbnail is the asset's preview image through DocumentDataExtractableAlteration.

API reference​

Download interfaces​

ClassDescription
DownloadableBase interface for assets that can be downloaded.
DocumentDownloadableDownloads documents. Override downloadDocumentFiles(DocumentDownloadOptions<?>, Path) to supply a StorageItem.
ImageDownloadableDownloads images. Override downloadImageFiles(ImageDownloadOptions, Path).
DamDocumentEntryViewMarker interface for a view model that Brightspot renders when converting a DocumentDownloadable.
DownloadOptions<F>Describes a download target. The type F determines the formats that the Download Options window offers.
DocumentFormatAbstract class for a document output format. Implement getFileExtensions(), getContentTypes(), and download(InputStream, Path, String).
ImageFormatAbstract class for an image output format. Implement getFileExtension(), getContentTypes(), and download(InputStream, File).

Extraction classes​

ClassDescription
DocumentDataExtractableInterface that opts an asset type in to extraction.
DocumentDataExtractorAbstract class for an extraction service.
DocumentThumbnailExtractorInterface for an extractor that can also generate a thumbnail. Textract uses it as an optional thumbnail source.
DocumentDataExtractableDataModification that stores the extracted text and thumbnail.
DocumentDataExtractorSettingsStores the list of configured extractors in the CMS settings.

The following table lists the methods to implement in a DocumentDataExtractor subclass.

MethodReturn typeDescription
getText(StorageItem)StringReturns the extracted text for the file.
shouldRun(String)booleanReturns true if the extractor is configured and supports the given content type.
runService(DocumentDataExtractableData, StorageItem)voidStores the text and thumbnail on the data object and saves it.
getSupportedFileTypes()Set<String>Returns the supported content types. An empty set means all types.

Pandoc classes​

ClassDescription
PandocDownloadableInterface for documents that Pandoc converts.
PandocDocumentDownloadOptionsDownload options that run Pandoc on a StorageItem.
PandocDocumentFormatAbstract class for Pandoc output formats. Implement toFormatParameter(), and optionally getAdditionalOptions().
PandocFromFormatSource formats, such as WORD_DOCUMENT, HTML, and RESTRUCTUREDTEXT.

ImageMagick classes​

ClassDescription
ImageMagickDownloadableInterface for images that ImageMagick converts.
ImageMagickDownloadOptionsDownload options that run ImageMagick on a StorageItem.
ImageMagickFormatAbstract class for ImageMagick output formats. Implement getFileExtension(), and optionally getAdditionalOptions().

Writing a custom extractor​

Extend DocumentDataExtractor. After you deploy the class, it appears in the Extractor Services list under DAM Document Data Extraction Settings.

1
import java.io.IOException;
2
import java.io.InputStream;
3
import java.nio.charset.StandardCharsets;
4
import java.util.Set;
5
6
import com.psddev.dam.DocumentDataExtractableData;
7
import com.psddev.dam.DocumentDataExtractor;
8
import com.psddev.dari.util.CompactSet;
9
import com.psddev.dari.util.IoUtils;
10
import com.psddev.dari.util.StorageItem;
11
12
public class PlainTextExtractor extends DocumentDataExtractor {
13
14
@Override
15
public Set<String> getSupportedFileTypes() {
16
return new CompactSet<>(Set.of("text/plain"));
17
}
18
19
@Override
20
public boolean shouldRun(String fileType) {
21
return getSupportedFileTypes().contains(fileType);
22
}
23
24
@Override
25
public String getText(StorageItem documentFile) {
26
try (InputStream data = documentFile.getData()) {
27
return IoUtils.toString(data, StandardCharsets.UTF_8);
28
29
} catch (IOException error) {
30
return null;
31
}
32
}
33
34
@Override
35
public void runService(DocumentDataExtractableData data, StorageItem documentFile) {
36
data.setText(getText(documentFile));
37
data.save();
38
}
39
}

Configuration​

Administrators set most options in the CMS. Operators can set the following values in the application's settings.

SettingDefaultDescription
dam/pandoc/executablepandocPath to the Pandoc executable.
dam/pandoc/processPathNoneDirectories to add to the path of the Pandoc process.
dam/pandoc/timeoutMillis5000Time in milliseconds before a Pandoc conversion fails.
dam/imagemagick/executableconvertPath to the ImageMagick executable.
dam/imagemagick/processPathNoneDirectories to add to the path of the ImageMagick process.
dam/imagemagick/timeoutMillis5000Time in milliseconds before an ImageMagick conversion fails.

Settings entered in the CMS take precedence over these values. For the CMS options, see Configuration.

Deprecated features​

  • SharedCollection is deprecated since 4.8. Use collections in the core platform.
  • The CloudConvertSettings modification is deprecated. Use CloudConvertDocumentDataExtractor in DAM Document Data Extraction Settings.

Was this page helpful?

This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.