A script to import every markdown document under a specified folder into Confluence. It handles inline images as well as code blocks. Also there is support for some custom markdown tags for use with commonly used Confluence macros.
Each file will be converted into HTML or Confluence storage markup when required. Then a page will be created or updated in the space. The hierarchy of the Confluence pages will mirror the folder structure under docs. Every folder has to have a markdown file under docs with the same name as the folder, to allow generating a corresponding page in the hierarchy. If a file is deleted, then running the tool will also remove the Confluence page. When a file is moved, then it takes about 24 hours for Confluence to rebuild the ancestor tree, so the change does not show up immediately.
Windows:
macOS:
Linux:
There are three options for authentication:
To generate an API key go to https://id.atlassian.com/manage/api-tokens.
For more information on Personal Access Tokens see https://confluence.atlassian.com/enterprise/using-personal-access-tokens-1026032365.html.
These can be set as command-line parameters; however, it is recommended that these instead should be set in a confluence.env file and passed into the docker run --env-file confluence.env ... command. The confluence.env file should then be excluded from checking into source control in order to keep the credentials secret. This also prevents the credentials from appearing in your shell history.
Additionally, you will need to know your organization name.
https://fawltytowers.atlassian.net/wiki/ then the organization name is fawltytowers.https://fawltytowers.mydomain.com/ then the organization name is fawltytowers.mydomain.com.If you are using password or API key based auth, the confluence.env file should look like:
CONFLUENCE_USERNAME=basil
CONFLUENCE_API_KEY=abc123
CONFLUENCE_ORGNAME=fawltytowers
If you are using Personal Access Token for auth, the confluence.env file should look like:
CONFLUENCE_PERSONAL_ACCESS_TOKEN=[...]
CONFLUENCE_ORGNAME=fawltytowers
Within Confluence, you will need to know a parent page ID under which to publish the pages that are uploaded. Finding this in the Confluence web UI can be a bit tricky; here's how you can do it:
https://fawltytowers.mydomain.com/pages/viewpreviousversions.action?pageId=1234567890pageId query parameter; in this example it is 1234567890The minimum accepted parameters are:
.env file (see above)/publish mount pointdocker run --rm \
--env-file confluence.env `# authentication parameters in confluence.env file` \
-v $(cwd):/publish `# publishes current working directory and subdirectories` \
tomkludy/md_to_conf \
Test-Space `# Confluence space key to publish to` \
-a 1234567890 `# ancestor page id`
| Parameter | Usage | |
|---|---|---|
| ‑t | ‑‑pat | Confluence personal access token if CONFLUENCE_PERSONAL_ACCESS_TOKEN is not set in the environment file. Either PAT or API key is Required. |
| ‑u | ‑‑username | Confluence username if CONFLUENCE_USERNAME is not set in the environment file. Required only if using API key. |
| ‑p | ‑‑apikey | Confluence password or API key if CONFLUENCE_API_KEY is not set in the environment file. Either PAT or API key is Required. |
| ‑o | ‑‑orgname | Required. Confluence organization name if CONFLUENCE_ORGNAME is not set in the environment file. If orgname contains a dot, it will be considered as the fully qualified domain name. |
| ‑a | ‑‑ancestor | Required. The id of the parent page under which every other page will be created or updated. |
| ‑‑note | Specifies a note to prepend on generated html pages. Useful to indicate to the user that the page is generated. | |
| ‑n | ‑‑nossl | If specified, will use HTTP instead of HTTPS. |
| ‑l | ‑‑loglevel | Set the log verbosity. Default: INFO |
Note: There are some additional undocumented options available, which may be complex to use from within docker. Use
docker run --rm tomkludy/md_to_conf -hto view a list of all available options.
Older versions of this tool (up to v1.0.4) included a -s / --simulate option. Unfortunately this option was buggy and not very useful since the output was hard to interpret. The option has been removed in v1.1.0. To avoid causing unexpected changes, if this command line option is specified, the tool will abort rather than proceed.
The recommended way to test changes before publishing is to target a personal space, typically ~username. Either create an empty page to use as the ancestor, or copy the doc tree from your public space. Publish your docs there before publishing to the public space. This will allow you to visually review all of the pages before continuing.
Note that some shells will expand ~username to /home/username, so make sure you quote the space name when passing it as a command line parameter.
The original markdown to HTML conversion is performed by the Python markdown library. Additionally, the page name is taken from the first line of each markdown file, usually assumed to be the title. In the case of this document, the page would be called: Documentation.
Standard markdown syntax for images and code blocks will be automatically converted. The images are uploaded as attachments and the references updated in the HTML. The code blocks will be converted to the Confluence Code Block macro and also supports syntax highlighting.
If present, what is between the doctoc anchor format:
<!-- START doctoc ...
...
... END doctoc -->
will be replaced by confluence "toc" macro leading to something like:
<h2>Table of Content</h2>
<p>
<ac:structured-macro ac:name="toc">
<ac:parameter ac:name="printable">true</ac:parameter>
<ac:parameter ac:name="style">disc</ac:parameter>
<ac:parameter ac:name="maxLevel">7</ac:parameter>
<ac:parameter ac:name="minLevel">1</ac:parameter>
<ac:parameter ac:name="type">list</ac:parameter>
<ac:parameter ac:name="outline">clear</ac:parameter>
<ac:parameter ac:name="include">.*</ac:parameter>
</ac:structured-macro>
</p>
Warning: Any blockquotes used will implement an information macro. This could potentially harm your formatting.
Block quotes in Markdown are rendered as information macros.
> This is an info

> Note: This is a note

> Warning: This is a warning

Alternatively, using a custom Markdown syntax also works:
~?This is an info.?~
~!This is a note.!~
~%This is a warning.%~
The intention is to be able to repeatedly publish updates to a Confluence page tree based on source files that change over time. For this reason the tool will attempt to minimize changes within Confluence. If an existing page's contents are updated or a page is moved to a different folder, the existing page is updated rather than recreated. This allows the comments attached to the page to be retained. It also allows the page history to be viewed in Confluence.
However if a page is renamed (by changing the first line in the file), it will be treated as a brand-new page and created anew. The old page will be moved to __ORPHAN__ and the change history and comments will be associated with the old page, not the new one.
Page contents will always be updated to match the source file if there are any discrepencies, so any manual edits to published pages will be overwritten when the page is next published. It's recommended to include a warning note in every generated page. This can be achieved using the -note command line parameter or by setting the CONFLUENCE_NOTE property in the environment file; for example:
CONFLUENCE_NOTE=This is a generated file. Any modifications to it will be lost upon next update. The source files are in the <a href="https://source-project-repo-url">project repository</a>.
All pages that are more than one level below the ancestor page must exist in the source folder, or else they will be considered orphaned and will be moved to a folder named __ORPHAN__.
For example, if Confluence contains a page structure:
Ancestor
+-- Page1
+-- Child1.1
+-- Child1.2
+-- Page2
+-- Child2.1
And the source folder contains a page structure:
Ancestor
+-- Page1
+-- Child1.1
Then the tool will update the Confluence page structure to:
Ancestor
+-- Page1
+-- Child1.1
+-- Page2
+-- Child2.1
+-- __ORPHAN__
+-- Child1.2
Note that Page2 is left alone. The assumption is that sibling pages to the published assets that may be completely unrelated to the published folder, while Child1.2 is folder or page that falls into the scope of the published folder. Its absence most likely indicates that a source file was deleted since the last time it was published, so the page is also likely to need to be deleted.
Note that this means that top-level pages will never be automatically orphaned once published, even if it was originally published by the tool and then the source file was deleted.
The __ORPHAN__ folder allows you to review pages before deletion. If you find the pages are safe to delete after reviewing, you can delete the __ORPHAN__ folder through the Confluence UI or API.
Content type
Image
Digest
Size
22.4 MB
Last updated
about 4 years ago
docker pull tomkludy/md_to_conf