Sign inSign up

tomkludy/md_to_conf

By tomkludy

•Updated about 4 years ago

Upload markdown to Confluence

Image
0

10K+

tomkludy/md_to_conf repository overview

⁠Documentation

⁠Markdown to Confluence Converter

A script to import every markdown document under a specified folder into Confluence. It handles inline images as well as code blocks. Also there is support for some custom markdown tags for use with commonly used Confluence macros.

Each file will be converted into HTML or Confluence storage markup when required. Then a page will be created or updated in the space. The hierarchy of the Confluence pages will mirror the folder structure under docs. Every folder has to have a markdown file under docs with the same name as the folder, to allow generating a corresponding page in the hierarchy. If a file is deleted, then running the tool will also remove the Confluence page. When a file is moved, then it takes about 24 hours for Confluence to rebuild the ancestor tree, so the change does not show up immediately.

⁠Use
⁠Prerequisites

Windows:

macOS:

Linux:

⁠Set up authentication

There are three options for authentication:

  1. Username + password
  2. Username + API key
  3. Personal Access Token

To generate an API key go to https://id.atlassian.com/manage/api-tokens⁠.

For more information on Personal Access Tokens see https://confluence.atlassian.com/enterprise/using-personal-access-tokens-1026032365.html⁠.

These can be set as command-line parameters; however, it is recommended that these instead should be set in a confluence.env file and passed into the docker run --env-file confluence.env ... command. The confluence.env file should then be excluded from checking into source control in order to keep the credentials secret. This also prevents the credentials from appearing in your shell history.

Additionally, you will need to know your organization name.

  • If you are using Confluence Cloud, you will need the organization name that is used in the subdomain. For example, if you normally access the URL: https://fawltytowers.atlassian.net/wiki/ then the organization name is fawltytowers.
  • If you are using Confluence On-Prem, you will need the Fully Qualified Domain Name of your server. For example, if you normally access the URL: https://fawltytowers.mydomain.com/ then the organization name is fawltytowers.mydomain.com.

If you are using password or API key based auth, the confluence.env file should look like:

CONFLUENCE_USERNAME=basil
CONFLUENCE_API_KEY=abc123
CONFLUENCE_ORGNAME=fawltytowers

If you are using Personal Access Token for auth, the confluence.env file should look like:

CONFLUENCE_PERSONAL_ACCESS_TOKEN=[...]
CONFLUENCE_ORGNAME=fawltytowers
⁠Additional requirements

Within Confluence, you will need to know a parent page ID under which to publish the pages that are uploaded. Finding this in the Confluence web UI can be a bit tricky; here's how you can do it:

  1. Navigate to the page that you want to be the root page in your browser
  2. Click the "three dots" menu near the upper right of the page
  3. Right-click and Copy Link (don't left-click) the "Page History" link; this should be something like https://fawltytowers.mydomain.com/pages/viewpreviousversions.action?pageId=1234567890
  4. Extract the pageId query parameter; in this example it is 1234567890
⁠Use
⁠Basic

The minimum accepted parameters are:

  • The authentication parameters, preferably contained within a .env file (see above)
  • The folder containing .md files to upload; this must be mapped into the container as a volume at the /publish mount point
  • The Confluence space key you wish to upload to
  • The ancestor page id, under which all files will be uploaded
docker run --rm \
    --env-file confluence.env `# authentication parameters in confluence.env file` \
    -v $(cwd):/publish        `# publishes current working directory and subdirectories` \
    tomkludy/md_to_conf \
    Test-Space                `# Confluence space key to publish to` \
    -a 1234567890             `# ancestor page id`
⁠Command line arguments
ParameterUsage
‑t‑‑patConfluence personal access token if CONFLUENCE_PERSONAL_ACCESS_TOKEN is not set in the environment file. Either PAT or API key is Required.
‑u‑‑usernameConfluence username if CONFLUENCE_USERNAME is not set in the environment file. Required only if using API key.
‑p‑‑apikeyConfluence password or API key if CONFLUENCE_API_KEY is not set in the environment file. Either PAT or API key is Required.
‑o‑‑orgnameRequired. Confluence organization name if CONFLUENCE_ORGNAME is not set in the environment file. If orgname contains a dot, it will be considered as the fully qualified domain name.
‑a‑‑ancestorRequired. The id of the parent page under which every other page will be created or updated.
‑‑noteSpecifies a note to prepend on generated html pages. Useful to indicate to the user that the page is generated.
‑n‑‑nosslIf specified, will use HTTP instead of HTTPS.
‑l‑‑loglevelSet the log verbosity. Default: INFO

Note: There are some additional undocumented options available, which may be complex to use from within docker. Use docker run --rm tomkludy/md_to_conf -h to view a list of all available options.

⁠Testing changes before publishing

Older versions of this tool (up to v1.0.4) included a -s / --simulate option. Unfortunately this option was buggy and not very useful since the output was hard to interpret. The option has been removed in v1.1.0. To avoid causing unexpected changes, if this command line option is specified, the tool will abort rather than proceed.

The recommended way to test changes before publishing is to target a personal space, typically ~username. Either create an empty page to use as the ancestor, or copy the doc tree from your public space. Publish your docs there before publishing to the public space. This will allow you to visually review all of the pages before continuing.

Note that some shells will expand ~username to /home/username, so make sure you quote the space name when passing it as a command line parameter.

⁠Markdown

The original markdown to HTML conversion is performed by the Python markdown library. Additionally, the page name is taken from the first line of each markdown file, usually assumed to be the title. In the case of this document, the page would be called: Documentation.

Standard markdown syntax for images and code blocks will be automatically converted. The images are uploaded as attachments and the references updated in the HTML. The code blocks will be converted to the Confluence Code Block macro and also supports syntax highlighting.

⁠Doctoc

If present, what is between the doctoc⁠ anchor format:

<!-- START doctoc ...
...
... END doctoc -->

will be replaced by confluence "toc" macro leading to something like:

<h2>Table of Content</h2>
<p>
    <ac:structured-macro ac:name="toc">
      <ac:parameter ac:name="printable">true</ac:parameter>
      <ac:parameter ac:name="style">disc</ac:parameter>
      <ac:parameter ac:name="maxLevel">7</ac:parameter>
      <ac:parameter ac:name="minLevel">1</ac:parameter>
      <ac:parameter ac:name="type">list</ac:parameter>
      <ac:parameter ac:name="outline">clear</ac:parameter>
      <ac:parameter ac:name="include">.*</ac:parameter>
    </ac:structured-macro>
    </p>
⁠Information, Note and Warning Macros

Warning: Any blockquotes used will implement an information macro. This could potentially harm your formatting.

Block quotes in Markdown are rendered as information macros.

> This is an info

macros

> Note: This is a note

macros

> Warning: This is a warning

macros

Alternatively, using a custom Markdown syntax also works:

~?This is an info.?~

~!This is a note.!~

~%This is a warning.%~
⁠Page history / comments

The intention is to be able to repeatedly publish updates to a Confluence page tree based on source files that change over time. For this reason the tool will attempt to minimize changes within Confluence. If an existing page's contents are updated or a page is moved to a different folder, the existing page is updated rather than recreated. This allows the comments attached to the page to be retained. It also allows the page history to be viewed in Confluence.

However if a page is renamed (by changing the first line in the file), it will be treated as a brand-new page and created anew. The old page will be moved to __ORPHAN__ and the change history and comments will be associated with the old page, not the new one.

Page contents will always be updated to match the source file if there are any discrepencies, so any manual edits to published pages will be overwritten when the page is next published. It's recommended to include a warning note in every generated page. This can be achieved using the -note command line parameter or by setting the CONFLUENCE_NOTE property in the environment file; for example:

CONFLUENCE_NOTE=This is a generated file. Any modifications to it will be lost upon next update. The source files are in the <a href="https://source-project-repo-url">project repository</a>.
⁠Orphan pages

All pages that are more than one level below the ancestor page must exist in the source folder, or else they will be considered orphaned and will be moved to a folder named __ORPHAN__.

For example, if Confluence contains a page structure:

Ancestor
+-- Page1
   +-- Child1.1
   +-- Child1.2
+-- Page2
   +-- Child2.1

And the source folder contains a page structure:

Ancestor
+-- Page1
   +-- Child1.1

Then the tool will update the Confluence page structure to:

Ancestor
+-- Page1
   +-- Child1.1
+-- Page2
   +-- Child2.1
+-- __ORPHAN__
   +-- Child1.2

Note that Page2 is left alone. The assumption is that sibling pages to the published assets that may be completely unrelated to the published folder, while Child1.2 is folder or page that falls into the scope of the published folder. Its absence most likely indicates that a source file was deleted since the last time it was published, so the page is also likely to need to be deleted.

Note that this means that top-level pages will never be automatically orphaned once published, even if it was originally published by the tool and then the source file was deleted.

The __ORPHAN__ folder allows you to review pages before deletion. If you find the pages are safe to delete after reviewing, you can delete the __ORPHAN__ folder through the Confluence UI or API.

Tag summary

Content type

Image

Digest

Size

22.4 MB

Last updated

about 4 years ago

docker pull tomkludy/md_to_conf