Skip to content

Add VGT Corpus to list of datasets #74

Description

@cleong110
@misc{dataset:herreweghe2015VGTCorpus,
 author = {{Van Herreweghe, Mieke and Vermeerbergen, Myriam and Demey, Eline and De Durpel, Hannes and Nyffels, Hilde and Verstraete, Sam}},
 keywords = {{Vlaamse Gebarentaal,Corpuslinguïstiek}},
 language = {{dut}},
 title = {{Het Corpus VGT. Een digitaal open access corpus van videos and annotaties van Vlaamse Gebarentaal, ontwikkeld aan de Universiteit Gent ism KU Leuven. <www.corpusvgt.be>}},
 url = {{http://www.corpusvgt.ugent.be/}},
 year = {{2015}}
}

Additional Checklist for datasets:

When adding a dataset, follow the following steps. This pull request provides an example:

  • Fork the repo
  • sync forks
  • git checkout master
  • git pull
  • New branch: dataset/something
  • Create a JSON along the lines of the schema below. e.g. FOO.json
  • Add the JSON to src/datasets. e.g. src/datasets/FOO.json
  • Add BibTex to src/references.bib.
    • prepend the citation key with dataset. e.g. dataset:sehyr2021asl
    • "language" field should not need "sign language". No need to say "American Sign Language", "American" will do.
    • Very concise "samples" field, the table does not have a lot of space to display it.
  • Commit/push the changes
  • Make a pull request!

Schema:

{
  "pub": {
    "name": string, # this gets used as the name of the dataset, e.g. "WLASL"
    "year": integer or null,
    "publication":string or null, # this matches a key in references.bib, e.g. "dataset:joshiISLTranslateDatasetTranslating2023"
    "url": string or null # URL to access it. e.g. "https://www.sign-lang.uni-hamburg.de/dgs-korpus/index.php/welcome.html"
  },
  "#loader": string or null, # the key you would use in the sign language datasets library. e.g. "dgs_corpus". Website will auto-link
  "#items": integer or null, # this is the number of unique signs in the column
  "#samples": string or null, # e.g. "1100 videos" or "8,257 Sentences"
  "#signers": integer or string or null, # number of unique signers
  "features": array of strings, ["feature1","feature2"], # I've seen things like "mouthing", "video:RGB", "pose:Kinect", "pose:OpenPose","text:Polish", "gloss:ASL", "writing:HamNoSys", etc.
  "language": string, # the Sign language or languages, e.g. "American" for American Sign Language (ASL)
  "license": string or null,
  "licenseUrl": string or null
}

Activity

cleong110 commented on Jun 24, 2024

@cleong110
ContributorAuthor

Issues I encountered documenting this:

  • Documentation is in Dutch, I don't speak Dutch, and the English version of the site is not up. Google Translate on website helped.
  • Couldn't find the vocabulary size anywhere on the website.
  • The website says "click here to download the detailed methodology for annotations", but the link leads to a template PDF.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions